Closed Bug 8607 Opened 27 years ago Closed 27 years ago

Optimize charset conversion for intl documents

Categories

(Core :: XML, defect, P3)

x86
Windows NT
defect

Tracking

()

RESOLVED FIXED

People

(Reporter: harishd, Assigned: nisheeth_mozilla)

References

Details

Here is the sample: 1 <?xml version="1.0" encoding="UTF-8"?> 2 <?xml-stylesheet href="xul.css" type="text/css"?> 3 <!DOCTYPE window> 4 <xul:window xmlns:html="http://www.w3.org/TR/REC-html40" 5 ..... In the above example the PI in line# 1 did not get handled in CWellFormedDTD.
Status: NEW → ASSIGNED
Target Milestone: M9
Accepting bug and setting milestone to M8...
Target Milestone: M9 → M8
Summary: The first PI encountered is ignored. → The XML Decl and Doc Type Decl need to get passed up from expat
Target Milestone: M8 → M9
The "PI in line 1" is actually an XML declaration, different from ordinary PIs. We need to create a new kind of token for it in nsExpatTokenizer and call nsXMLContentSink::AddXMLDecl when we see that token inside nsWellFormedDTD. This work is too much to put into M8. Changing milestone to M9...
Whiteboard: Start of fix in home computer's moztemp directory.
Updating status whiteboard...
*** Bug 6046 has been marked as a duplicate of this bug. ***
Whiteboard: Start of fix in home computer's moztemp directory. → Need to work with Harish on genericizing token types.
Adding rickg to the cc list. He is planning to make similar changes to the HTML tokenizer and DTD once I've put in the XML changes.
Whiteboard: Need to work with Harish on genericizing token types.
James, I am currently using the default handler for accessing the XML decl and DOCTYPE decl. It would be great to have explicit callbacks for both these declarations like we discussed over email earlier. Please let us know when you plan to add these callbacks to expat. Thanks.
Why do you need the XML decl? Applications shouldn't care whether there's an XML decl or not. When writing out XML, you should generate an XML decl that is correct for what you are writing out (eg specifying the encoding that you are using to output the XML). This isn't in the DOM Level 1; it's not in SAX; it's not needed by XSLT.
Rick, I've added the aMode parameter to all the classes that implement nsIContentSink::AddDocTypeDecl and checked my changes in. To provide context for this bug report, earlier AddDocTypeDecl was a method on nsIXMLContentSink, but, as part of the work for this bug, the method has now moved to nsIContentSink. The parser is going to use the new aMode parameter to inform the content sink which mode to process the parser nodes in: HTML 4.0 strict, Nav quirks, etc.
I've added callbacks to expat to pass the DOCTYPE decl.
James, I agree with you about not adding an XML decl callback. I had thought that the application needs the XML decl for two reasons: a) to get at the encoding attribute, b) to pass up the XML decl to the XML content sink which had a method that consumed the XML decl. I realize now that both the reasons are invalid: a) Expat reports the encoding via a separate callback, b) Vidur had added the XML decl method on the XML content sink back when there had been discussion in the DOM group about exposing the XML decl via the DOM. I spoke to him recently and he, like you, doesn't think we need to keep the method around. So, I'm going to update expat to your latest test release that has DOCTYPY start and end callbacks once the tree opens after the current Necko stability push.
I've checked in the latest version of expat provided by James and added code in the XML tokenizer and the DTD to propagate the DOCTYPE decl to the sink. James, it seems like I was mistakenly thinking that the unknown encoding handler will get called when an unsupported encoding attribute is specified on the XML decl. Is there some other way that I can get at the encoding attribute? The internationalization team need to know the encoding so that they can hook in a charset converter.
I thought you were converting the input data to UCS2 before calling expat. nsExpatTokenizer calls XML_ParserCreat() with an argument of "UTF-16", which tells expat to ignore the encoding declaration and use UTF-16. If you're converting the input data to UCS2 before calling expat, then you've got to know the encoding before you call expat, so why does expat need to tell you the encoding?
Status: ASSIGNED → RESOLVED
Closed: 27 years ago
Resolution: --- → FIXED
Summary: The XML Decl and Doc Type Decl need to get passed up from expat → The Doc Type Decl needs to get passed up from expat
Adjusting summary. I just discussed charset conversion with Frank Tang, an internationalization engineer. We decided that we will sniff the encoding and convert the incoming data to UCS2 before the data gets passed to expat. So, expat will always see UCS2. Till now, we converted the incoming data to UCS2 without sniffing the encoding (we assumed that the encoding of the incoming data was UTF-8) and passed on the UCS2 data to expat. The expectation was that if the encoding was non-UTF-8 and our guess was wrong, we would re-load the document and convert the incoming data to UCS2 using the specified encoding. This is why we needed the encoding callback from expat. Now that we'll determine the encoding before expat sees the data, we don't need the callback any more. CCing Frank Tang and marking this bug fixed.
The approach you now have in mind should work. Note that you should only sniff the encoding if the content-type is application/xml with no charset parameter. If the content-type is text/xml with a charset or application/xml with a charset parameter, then RFC 2376 requires you to use the specified charset parameter. If the content-type is text/xml with no charset parameter, RFC 2376 requires you to use us-ascii as the encoding. If you know the encoding from the content-type and the encoding is one that expat can handle internally (utf-8, utf-16, us-ascii, iso-8859-1), then it is much more efficient just to pass the data unconverted to expat, tell expat what the encoding is in XML_ParserCreate and let expat do the conversion. However, the approach you have in mind is less efficient that it need be. An approach more like what you originally had in mind would be more efficient: - pass the incoming data without conversion to expat - the encoding parameter passed to XML_ParserCreate should be the encoding determined by the content-type per RFC 2376 unless it is application/xml without a charset in which case it should be null - set an UnknownEncodingHandler; if that gets called then stop the parse; reload the document, this time converting the data into UTF-16 - better yet, use expat's unknown encoding handling machinery; for example, if you have a single byte encoding, the unknown encoding handler can pass expat a table that maps the encoding into Unicode, and then expat can do the conversion as part of the encoding process; only reload the document when you get an encoding of a type that expat's unknown encoding handling machinery cannot handle (ie a stateful encoding such as ISO-2022-JP)
QA Contact: chrisd → janc
Jan: This looks like a whitebox issue for verification. Please take a look. Thanks
James' comments deal with optimizing the parsing of international documents. I've created a new bug (bug 12375) with James' last comment pasted on it and assigned it to Frank Tang. We can track how James' suggestions get implemented on that bug.
Summary: The Doc Type Decl needs to get passed up from expat → Optimize charset conversion for intl documents
You need to log in before you can comment on or make changes to this bug.