Closed
Bug 8607
Opened 27 years ago
Closed 27 years ago
Optimize charset conversion for intl documents
Categories
(Core :: XML, defect, P3)
Tracking
()
RESOLVED
FIXED
M9
People
(Reporter: harishd, Assigned: nisheeth_mozilla)
References
Details
Here is the sample:
1 <?xml version="1.0" encoding="UTF-8"?>
2 <?xml-stylesheet href="xul.css" type="text/css"?>
3 <!DOCTYPE window>
4 <xul:window xmlns:html="http://www.w3.org/TR/REC-html40"
5 .....
In the above example the PI in line# 1 did not get handled in CWellFormedDTD.
| Assignee | ||
Updated•27 years ago
|
Status: NEW → ASSIGNED
Target Milestone: M9
| Assignee | ||
Comment 1•27 years ago
|
||
Accepting bug and setting milestone to M8...
| Assignee | ||
Updated•27 years ago
|
Target Milestone: M9 → M8
| Assignee | ||
Updated•27 years ago
|
Summary: The first PI encountered is ignored. → The XML Decl and Doc Type Decl need to get passed up from expat
Target Milestone: M8 → M9
| Assignee | ||
Comment 2•27 years ago
|
||
The "PI in line 1" is actually an XML declaration, different from ordinary PIs.
We need to create a new kind of token for it in nsExpatTokenizer and call
nsXMLContentSink::AddXMLDecl when we see that token inside nsWellFormedDTD.
This work is too much to put into M8. Changing milestone to M9...
| Assignee | ||
Updated•27 years ago
|
Whiteboard: Start of fix in home computer's moztemp directory.
| Assignee | ||
Comment 3•27 years ago
|
||
Updating status whiteboard...
| Assignee | ||
Updated•27 years ago
|
Whiteboard: Start of fix in home computer's moztemp directory. → Need to work with Harish on genericizing token types.
| Assignee | ||
Comment 5•27 years ago
|
||
Adding rickg to the cc list. He is planning to make similar changes to the HTML
tokenizer and DTD once I've put in the XML changes.
| Assignee | ||
Updated•27 years ago
|
Whiteboard: Need to work with Harish on genericizing token types.
| Assignee | ||
Comment 6•27 years ago
|
||
James, I am currently using the default handler for accessing the XML decl and
DOCTYPE decl. It would be great to have explicit callbacks for both these
declarations like we discussed over email earlier. Please let us know when you
plan to add these callbacks to expat. Thanks.
Comment 7•27 years ago
|
||
Why do you need the XML decl? Applications shouldn't care whether there's an XML
decl or not. When writing out XML, you should generate an XML decl that is
correct for what you are writing out (eg specifying the encoding that you are
using to output the XML).
This isn't in the DOM Level 1; it's not in SAX; it's not needed by XSLT.
| Assignee | ||
Comment 8•27 years ago
|
||
Rick, I've added the aMode parameter to all the classes that implement
nsIContentSink::AddDocTypeDecl and checked my changes in.
To provide context for this bug report, earlier AddDocTypeDecl was a method on
nsIXMLContentSink, but, as part of the work for this bug, the method has now
moved to nsIContentSink. The parser is going to use the new aMode parameter to
inform the content sink which mode to process the parser nodes in: HTML 4.0
strict, Nav quirks, etc.
Comment 9•27 years ago
|
||
I've added callbacks to expat to pass the DOCTYPE decl.
| Assignee | ||
Comment 10•27 years ago
|
||
James, I agree with you about not adding an XML decl callback. I had thought
that the application needs the XML decl for two reasons: a) to get at the
encoding attribute, b) to pass up the XML decl to the XML content sink which had
a method that consumed the XML decl.
I realize now that both the reasons are invalid: a) Expat reports the encoding
via a separate callback, b) Vidur had added the XML decl method on the XML
content sink back when there had been discussion in the DOM group about exposing
the XML decl via the DOM. I spoke to him recently and he, like you, doesn't
think we need to keep the method around.
So, I'm going to update expat to your latest test release that has DOCTYPY start
and end callbacks once the tree opens after the current Necko stability push.
| Assignee | ||
Comment 11•27 years ago
|
||
I've checked in the latest version of expat provided by James and added code in
the XML tokenizer and the DTD to propagate the DOCTYPE decl to the sink.
James, it seems like I was mistakenly thinking that the unknown encoding handler
will get called when an unsupported encoding attribute is specified on the XML
decl. Is there some other way that I can get at the encoding attribute? The
internationalization team need to know the encoding so that they can hook in a
charset converter.
Comment 12•27 years ago
|
||
I thought you were converting the input data to UCS2 before calling expat.
nsExpatTokenizer calls XML_ParserCreat() with an argument of "UTF-16", which
tells expat to ignore the encoding declaration and use UTF-16. If you're
converting the input data to UCS2 before calling expat, then you've got to know
the encoding before you call expat, so why does expat need to tell you the
encoding?
| Assignee | ||
Updated•27 years ago
|
Status: ASSIGNED → RESOLVED
Closed: 27 years ago
Resolution: --- → FIXED
Summary: The XML Decl and Doc Type Decl need to get passed up from expat → The Doc Type Decl needs to get passed up from expat
| Assignee | ||
Comment 13•27 years ago
|
||
Adjusting summary.
I just discussed charset conversion with Frank Tang, an internationalization
engineer. We decided that we will sniff the encoding and convert the incoming
data to UCS2 before the data gets passed to expat. So, expat will always see
UCS2. Till now, we converted the incoming data to UCS2 without sniffing the
encoding (we assumed that the encoding of the incoming data was UTF-8) and
passed on the UCS2 data to expat. The expectation was that if the encoding was
non-UTF-8 and our guess was wrong, we would re-load the document and convert
the incoming data to UCS2 using the specified encoding. This is why we needed
the encoding callback from expat. Now that we'll determine the encoding before
expat sees the data, we don't need the callback any more.
CCing Frank Tang and marking this bug fixed.
Comment 14•27 years ago
|
||
The approach you now have in mind should work. Note that you should only sniff
the encoding if the content-type is application/xml with no charset parameter.
If the content-type is text/xml with a charset or application/xml with a charset
parameter, then RFC 2376 requires you to use the specified charset parameter.
If the content-type is text/xml with no charset parameter, RFC 2376 requires you
to use us-ascii as the encoding.
If you know the encoding from the content-type and the encoding is one that
expat can handle internally (utf-8, utf-16, us-ascii, iso-8859-1), then it is
much more efficient just to pass the data unconverted to expat, tell expat what
the encoding is in XML_ParserCreate and let expat do the conversion.
However, the approach you have in mind is less efficient that it need be. An
approach more like what you originally had in mind would be more efficient:
- pass the incoming data without conversion to expat
- the encoding parameter passed to XML_ParserCreate should be the encoding
determined by the content-type per RFC 2376 unless it is application/xml without
a charset in which case it should be null
- set an UnknownEncodingHandler; if that gets called then stop the parse; reload
the document, this time converting the data into UTF-16
- better yet, use expat's unknown encoding handling machinery; for example, if
you have a single byte encoding, the unknown encoding handler can pass expat a
table that maps the encoding into Unicode, and then expat can do the conversion
as part of the encoding process; only reload the document when you get an
encoding of a type that expat's unknown encoding handling machinery cannot
handle (ie a stateful encoding such as ISO-2022-JP)
Updated•27 years ago
|
QA Contact: chrisd → janc
Comment 15•27 years ago
|
||
Jan: This looks like a whitebox issue for verification. Please take a look.
Thanks
| Assignee | ||
Comment 16•27 years ago
|
||
James' comments deal with optimizing the parsing of international documents.
I've created a new bug (bug 12375) with James' last comment pasted on it and
assigned it to Frank Tang. We can track how James' suggestions get implemented
on that bug.
Summary: The Doc Type Decl needs to get passed up from expat → Optimize charset conversion for intl documents
You need to log in
before you can comment on or make changes to this bug.
Description
•