Open
Bug 1176848
Opened 11 years ago
Updated 3 years ago
I Can't read some email contents, It seems ISO-2022-KR encoding problem.
Categories
(MailNews Core :: MIME, defect)
Tracking
(Not tracked)
NEW
People
(Reporter: songkheechan, Unassigned)
References
Details
Attachments
(1 file)
|
6.44 KB,
message/rfc822
|
Details |
User Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:38.0) Gecko/20100101 Firefox/38.0
Build ID: 20150525141253
Steps to reproduce:
I was update Thunderbird to 38 from 31.7.0., in these days.
Actual results:
But I can,t read some email contents(not a subject, or attached file, only contents). It only shows '?'.
Then, I changed character encoding at view > character encoding, to Unicode(U), Korean(K), and any other characters, but I can't read, neither.
Also, in attachment field, there are some unknown file(s) was added, named like 'part1.1' 'part1.2' etc, with orignal file(s) which sender attached.
(But sample eml didn't have attached files.)
Expected results:
It must be shown correctly.
I think it comes from ISO-2022-KR encoding feature removed at version 38.
Because the source of mail which I can't read, contains below code in it.
Content-Type: text/plain; charset="ISO-2022-KR"
Please, help me.
(In reply to Kheechan from comment #0)
> Created attachment 8625444 [details]
> test with MIME.eml
You received this from <helpdesk@kr.ibm.com> ?
They really shouldn't do this
See Also: → https://support.mozilla.org/bs/questions/1068366
I don't know what you are mean, 'They really shouldn't do this'.
The sample eml is just a sample.
And ISO-2022-KR format is still using at some other mail clients, I think.
Comment 4•11 years ago
|
||
(In reply to Kheechan from comment #3)
> And ISO-2022-KR format is still using at some other mail clients, I think.
Do you have evidence on which ones?
Updated•11 years ago
|
Status: UNCONFIRMED → NEW
Ever confirmed: true
Comment 5•11 years ago
|
||
According to telemetry, before removal, ISO-2022-KR was used in 0.015% of Thunderbird sessions. It was the second-highest usage (after IBM852) for the encodings that were removed in Thunderbird 38.
(Bug 1155539 comment 2 for the archived telemetry readouts.)
Comparative data for EUC-KR was not collected.
Maybe, in fact, sample eml above composed by IBM Lotus Notes in korea.
Only some email composed from Notes, couldn't read.
Those email's source contains below, all.
Content-Type: text/plain; charset="ISO-2022-KR"
But some email composed from Notes equally, which can read, contains below in their source.
Content-type: text/plain; charset=UTF-8
Content-transfer-encoding: base64
So, I can think and tell to you,dear,
'the problem comes from ISO-2022-KR removal'.
Thanks.
Addendum;
If restoring ISO-2022-KR encoding is impossible,
then why don't you make it to install as a additional feature,
to whom want to use that encoding(like me).
Comment 8•11 years ago
|
||
(In reply to Kheechan from comment #6)
> So, I can think and tell to you,dear,
> 'the problem comes from ISO-2022-KR removal'.
Of course. I'm just trying to understand the extent of the impact of the problem.
(In reply to Kheechan from comment #7)
> If restoring ISO-2022-KR encoding is impossible,
> then why don't you make it to install as a additional feature,
> to whom want to use that encoding(like me).
It's not impossible to restore the code and it's probably less trouble to restore it unconditionally in Thunderbird than to offer it as a separately installable feature.
For background:
On the Web, ISO-2022-KR is a cross-site scripting hazard. Since ISO-2022-KR is not required for Web compatibility, it was removed from the Encoding Standard (https://encoding.spec.whatwg.org/ ; deals with encodings needed for supporting *Web content* rather than email) and from Firefox.
It is possible for Thunderbird to support an encoding for email even when support for the Web does not exist in Firefox. UTF-7 is currently such an encoding. It was decided in bug 1025886 not pursue that route for ISO-2022-KR on the grounds of evidence available from the records from the time when ISO-2022-KR support was added. Knowledge about the behavior of IBM Notes was not available at that time--it's clearly new information and Thunderbird developers can now decide if the new information warrants the resurrection of ISO-2022-KR.
It's worth noting that apart from the usual XSS hazards, the specific implementations of the ISO-2022-* series in the Mozilla code base have had at least one CVE-worthy security bug, which is why keeping the ISO-2022-KR code around in Thunderbird didn't look attractive in the absence of strong evidence of the necessity of keeping that code around.
(In reply to Kheechan from comment #3)
> I don't know what you are mean, 'They really shouldn't do this'.
> The sample eml is just a sample.
I mean <kr.ibm.com> should use utf-8, regardless of this bug.
Maybe it's worth to contact them?
(In reply to Henri Sivonen (:hsivonen) from comment #4)
> (In reply to Kheechan from comment #3)
> > And ISO-2022-KR format is still using at some other mail clients, I think.
>
> Do you have evidence on which ones?
Kheechan, do you know some?
| Reporter | ||
Comment 10•11 years ago
|
||
(In reply to j.j. from comment #9)
> (In reply to Kheechan from comment #3)
> > I don't know what you are mean, 'They really shouldn't do this'.
> > The sample eml is just a sample.
>
> I mean <kr.ibm.com> should use utf-8, regardless of this bug.
> Maybe it's worth to contact them?
>
I think the IBM Notes still using ISO-2022-kr format,
and I got many email comes from IBMer everyday.
I want to use Thunderbird for a long time.
I dont't want to change to other mail client application.
So, It is worth for me.(maybe, in korea, there are some other people who are in same situation like me, I thought)
I heard the IBM Notes can change mail format when compose mail, between Rich/Text and MIME type.
Composed email selected with Rich Text doesn't have problem, and their source contains 'Content-type: text/plain; charset=UTF-8'.
But, the email selected with MIME type, make this problem, and their source contains 'Content-Type: text/plain; charset="ISO-2022-KR"'.
General Notes users(maybe IBMer)don't know how to change or select email format to Rich/tesxt or MIME, they just 'use' indifferently.
And I can't require to them to change mail format to Rich/text, every time, manually, when I got unreadable email.
>
> (In reply to Henri Sivonen (:hsivonen) from comment #4)
> > (In reply to Kheechan from comment #3)
> > > And ISO-2022-KR format is still using at some other mail clients, I think.
> >
> > Do you have evidence on which ones?
>
> Kheechan, do you know some?
the IBM Notes, the only one that I found.
Comment 11•11 years ago
|
||
It is impractical, if not impossible, to support every charset. Unfortunately, it is also fairly difficult to work out a list of charsets to support. In trying to work out the list of charsets currently supported, I did collect information from at least 3 different sources of information, and while the representative nature of those datasets could be called into question, all of them did firmly plant ISO-2022-KR very much in the long tail of encodings. This was only reinforced by the Telemetry data indicating relatively little usage of ISO-2022-KR.
ISO-2022-KR is a mode-switching charset, which makes it all the more dangerous. Its special-casing in the Encoding spec (mapping to the replacement) makes it more annoying, but not impossible, to support in Thunderbird. For me at least, it needs a particularly compelling reason to keep it around--the rationale so far would probably have led me to WONTFIX this bug on the spot were it not that it used to be implemented.
(In reply to Kheechan from comment #7)
> If restoring ISO-2022-KR encoding is impossible,
> then why don't you make it to install as a additional feature,
> to whom want to use that encoding(like me).
It might be possible to make ISO-2022-KR available in an addon, but I've talked about changing the internal implementation of stuff in a way that would make that addon impossible to write in future versions.
I did a web search for ISO-2022-KR, and most of the results that turned up were about ISO-2022-JP instead. Of the few results that weren't, the linked support request and the original addition bug in Mozilla happened to be the first two. Most discussions I've seen on emailing lists indicate that the Korean variant of ISO-2022 fell out of favor rapidly (pre-2000!) (and the Chinese variant really appears to be unused).
I'd want to see firmer evidence that ISO-2022-KR is used outside of a misconfigured or minor-use application (in which case, evangelism would be greatly preferable) before consenting to reincluding a decoder. There is precedent--nearly all present use of x-mac-croatian came from recent versions of Thunderbird and is likely due to the presentation of that charset as "Croatian" in the encodings menu.
| Reporter | ||
Comment 12•11 years ago
|
||
Thank you for you positvive answer.
I have a question about this issue.
It seems Microsoft Outlook, MAC OS X mail, still support this ISO encoding,
because I can read those email on those email client softwares.
Is it right? are they fearless? or they are using another encoding methods?
Comment 13•11 years ago
|
||
(In reply to Kheechan from comment #12)
> Thank you for you positvive answer.
>
> I have a question about this issue.
> It seems Microsoft Outlook, MAC OS X mail, still support this ISO encoding,
> because I can read those email on those email client softwares.
> Is it right? are they fearless? or they are using another encoding methods?
Charset decoding is basically black magic. There is no real official list of charsets (the IANA list both misses important charsets and includes charsets that see no use anywhere, e.g., UTF-1), and this is compounding by many charsets having (sometimes incompatible) "extensions" not found in their official specification. On top of that, software ends up mislabelling charsets outright (definitely at least 0.1%, probably on the order of 1-6%, depending on what you consider mislabeled). It's only in 2012 that someone tried building an "official" list for the web, and it's still not clear how that list needs to be modified to account for email.
What happens in practice is that many clients defer charset detection/decoding to some library, and the quality and correctness of these libraries or of the use of these libraries (e.g., vanilla ICU is most likely going to screw over your users) is very much in wide variation. Particularly for email, the default modus operandi for most maintainers is extreme conservatism, which can be problematic for software with decades-long pedigrees (well, we needed this in 1990, I don't know if we still need it today). Gecko's drive to remove obsolete charsets is rather unprecedented (beyond some half-hearted efforts to eradicate EBCDIC), so we are treading unknown waters here.
Updated•11 years ago
|
Component: Untriaged → MIME
Product: Thunderbird → MailNews Core
Updated•3 years ago
|
Severity: normal → S3
You need to log in
before you can comment on or make changes to this bug.
Description
•