← Prev in month ← Prev in thread

Parsing bad HTML

From
Paul M <>
To
Date
2008-11-13T20:22:22Z
ID
<>
Thread
Parsing bad HTML
I use tidy to clean up  bad html docs. It does a pretty good job of converting html => strict xthml

However, the following is a bit too much

<p>
<sub>123</sub>4567<eight<img src="file.gif" alt="<b>hello</b>">
</p>

The problem is with 7<eight. Stray < and > seem to make tidy choke. What is the best method of handling this? I am leaning toward perl and regexp, but am hoping to avoid this. Maybe a Java solution? And tidy solutions?

-thanks

← Prev in month ← Prev in thread