清除Word转html的垃圾代码

最新推荐文章于 2021-12-08 16:15:59 发布

linshutao

最新推荐文章于 2021-12-08 16:15:59 发布

阅读量4.5k

点赞数

分类专栏： Java WEB开发文章标签： html attributes tags character microsoft class

Java 同时被 2 个专栏收录

51 篇文章 0 订阅

订阅专栏

WEB开发

43 篇文章 0 订阅

订阅专栏

Clean Word HTML using Regular Expressions

2005年11月23日 15:40:36 (GMT Standard Time, UTC+00:00) ( .Net General )

Introduction

I've spent a long time trying many different approaches at getting rid of MS Word HTML, when importing or pasting text into my content management system, with very mixed success. Previous efforts involved using the MSHTML Element Dom but this was slow and difficult to implement. i think i've finally found a satisfactory and fast solution using only regular expressions. Please feel free to use it in your applications, and post any improvements you may find.

The Code

/// <summary>
/// Removes all FONT and SPAN tags, and all Class and Style attributes.
/// Designed to get rid of non-standard Microsoft Word HTML tags.
/// </summary>
private string CleanHtml(string html)
{ 
    // start by completely removing all unwanted tags 
    html = Regex.Replace(html, @"<[/]?(font|span|xml|del|ins|[ovwxp]:\w+)[^>]*?>", "", RegexOptions.IgnoreCase); 
    // then run another pass over the html (twice), removing unwanted attributes 
    html = Regex.Replace(html, @"<([^>]*)(?:class|lang|style|size|face|[ovwxp]:\w+)=(?:'[^']*'|""[^""]*""|[^\s>]+)([^>]*)>","<$1$2>", RegexOptions.IgnoreCase); 
    html = Regex.Replace(html, @"<([^>]*)(?:class|lang|style|size|face|[ovwxp]:\w+)=(?:'[^']*'|""[^""]*""|[^\s>]+)([^>]*)>","<$1$2>", RegexOptions.IgnoreCase); 
    return html;
}

Samples of non-standard Microsoft Word HTML

<SPAN lang=EN-IE style="mso-ansi-language: EN-IE">
<p class="MSO Normal">
<UL style="MARGIN-TOP: 0cm" type=circle>
<o:p> </o:p>
<li class=MsoNormal style='mso-list:l3 level1 lfo3;tab-stops:list 36.0pt'>

Explanation of Regular Expressions

I've spent a good deal of time examining the problematic tags that MS Word inserts in its HTML, some examples are shown above. The above code is based on a few requirements for my CMS:

remove all FONT and SPAN tags, because all the content in my CMS is done through style-sheets.
remove all CLASS and STYLE tags because they mean nothing outside of the original word document
remove all namespace tags and attributes like <o:p> and < ... v:shape ... >

The first regular expression removes unwanted tags, and is broken down as follows:

<[/]?(font|span|xml|del|ins|[ovwxp]:\w+)[^>]*?>

match an open tag character <
and optionally match a close tag sequence </ (because we also want to remove the closing tags)
match any of the list of unwanted tags: font,span,xml,del,ins
a pattern is given to match any of the namespace tags, anything beginning with o,v,w,x,p, followed by a : followed by another word
match any attributes as far as the closing tag character >
the replace string for this regex is "", which will completely remove the instances of any matching tags.
note that we are not removing anything between the tags, just the tags themselves

The second regular expression removes unwanted attributes, and is broken down as follows:

<([^>]*)(?:class|lang|style|size|face|[ovwxp]:\w+)=(?:'[^']*'|""[^""]*""|[^\s>]+)([^>]*)>

match an open tag character <
capture any text before the unwanted attribute (This is $1 in the replace expression)
match (but don't capture) any of the unwanted attributes: class, lang, style, size, face, o:p, v:shape etc.
there should always be an = character after the attribute name
match the value of the attribute by identifying the delimiters. these can be single quotes, or double quotes, or no quotes at all.
for single quotes, the pattern is: ' followed by anything but a ' followed by a '
similarly for double quotes.
for a non-delimited attribute value, i specify the pattern as anything except the closing tag character >
lastly, capture whatever comes after the unwanted attribute in ([^>]*)
the replacement string <$1$2> reconstructs the tag without the unwanted attribute found in the middle.
note: this only removes one occurence of an unwanted attribute, this is why i run the same regex twice. For example, take the html fragment: <p class="MSO Normal" style="Margin-TOP:3em">
the regex will only remove one of these attributes. Running the regex twice will remove the second one. I can't think of any reasonable cases where it would need to be run more than that.

Suggestions!

If you have any suggestions or improvments, please post them here as comments.
Thanks :)

p.s. thanks to BinBin for the fix to preserve attributes like 'align=center'.

linshutao

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
清除Word转html的垃圾代码

Clean Word HTML using Regular Expressions2005年11月23日 15:40:36 (GMT Standard Time, UTC+00:00) ( .Net General )IntroductionI've spent a long time trying many different approaches at getting ri
复制链接

扫一扫