Parsing bad HTML

[Date Prev] | [Thread Prev] | [Thread Next] | [Date Next] -- [Date Index] | [Thread Index]

Parsing bad HTML

From: Paul M <pjmaip@yahoo.com>
To: xml-dev@lists.xml.org
Date: Thu, 13 Nov 2008 12:22:22 -0800 (PST)

I use tidy to clean up bad html docs. It does a pretty good job of converting html => strict xthml

However, the following is a bit too much


1234567<eight<img src="javascript:void(0);" alt="hello">


The problem is with 7<eight. Stray < and > seem to make tidy choke. What is the best method of handling this? I am leaning toward perl and regexp, but am hoping to avoid this. Maybe a Java solution? And tidy solutions?

-thanks

Follow-Ups:
- Re: [xml-dev] Parsing bad HTML
 - From: Henri Sivonen <hsivonen@iki.fi>
- Re: [xml-dev] Parsing bad HTML
 - From: "Steven J. DeRose" <sderose@acm.org>
- Re: [xml-dev] Parsing bad HTML
 - From: David Carlisle <davidc@nag.co.uk>
- Re: [xml-dev] Parsing bad HTML
 - From: COUTHURES Alain <alain.couthures@agencexml.com>

[Date Prev] | [Thread Prev] | [Thread Next] | [Date Next] -- [Date Index] | [Thread Index]