Save as rfc 2557 MHTML; complete webpage in one file
Categories
(Core :: DOM: Serializers, enhancement, P3)
Tracking
()
People
(Reporter: sidr, Assigned: jkt)
References
(Depends on 2 open bugs, Blocks 1 open bug, )
Details
(4 keywords)
Attachments
(9 files, 4 obsolete files)
|
74.18 KB,
text/plain
|
Details | |
|
64.57 KB,
text/plain
|
Details | |
|
40.43 KB,
patch
|
Details | Diff | Splinter Review | |
|
161.85 KB,
text/plain
|
Details | |
|
215.39 KB,
patch
|
Details | Diff | Splinter Review | |
|
48 bytes,
text/x-phabricator-request
|
Details | Review | |
|
341.81 KB,
image/png
|
Details | |
|
2.35 MB,
image/png
|
Details | |
|
48 bytes,
text/x-phabricator-request
|
Details | Review |
Comment 1•26 years ago
|
||
Comment 3•26 years ago
|
||
Comment 5•26 years ago
|
||
| Reporter | ||
Comment 6•26 years ago
|
||
Comment 8•25 years ago
|
||
Comment 10•25 years ago
|
||
Comment 11•25 years ago
|
||
Comment 12•25 years ago
|
||
Updated•25 years ago
|
Comment 13•25 years ago
|
||
Comment 14•24 years ago
|
||
Comment 15•24 years ago
|
||
Comment 16•24 years ago
|
||
Comment 17•24 years ago
|
||
Comment 18•24 years ago
|
||
Comment 19•24 years ago
|
||
Comment 20•24 years ago
|
||
Comment 21•24 years ago
|
||
Comment 22•24 years ago
|
||
Comment 23•24 years ago
|
||
Comment 24•24 years ago
|
||
Comment 25•24 years ago
|
||
Comment 26•24 years ago
|
||
Comment 27•24 years ago
|
||
Comment 28•24 years ago
|
||
Comment 29•24 years ago
|
||
Comment 30•24 years ago
|
||
Comment 31•24 years ago
|
||
Comment 32•24 years ago
|
||
Comment 33•24 years ago
|
||
Comment 34•24 years ago
|
||
Comment 35•24 years ago
|
||
Comment 36•24 years ago
|
||
Comment 37•24 years ago
|
||
Comment 38•24 years ago
|
||
Comment 39•24 years ago
|
||
Comment 40•24 years ago
|
||
Comment 41•24 years ago
|
||
Comment 42•24 years ago
|
||
Comment 43•24 years ago
|
||
Comment 44•24 years ago
|
||
Comment 45•24 years ago
|
||
Comment 46•24 years ago
|
||
Comment 47•24 years ago
|
||
Comment 48•24 years ago
|
||
Comment 49•24 years ago
|
||
Comment 50•24 years ago
|
||
Updated•23 years ago
|
Comment 51•23 years ago
|
||
Comment 52•23 years ago
|
||
Comment 53•23 years ago
|
||
Updated•23 years ago
|
Comment 54•23 years ago
|
||
Comment 55•23 years ago
|
||
Comment 56•23 years ago
|
||
Comment 57•23 years ago
|
||
Comment 58•23 years ago
|
||
Comment 59•23 years ago
|
||
Comment 60•23 years ago
|
||
Comment 61•23 years ago
|
||
Comment 62•23 years ago
|
||
Comment 63•23 years ago
|
||
Comment 64•23 years ago
|
||
Comment 65•23 years ago
|
||
Comment 66•23 years ago
|
||
Comment 67•23 years ago
|
||
Comment 68•23 years ago
|
||
Comment 69•23 years ago
|
||
Comment 70•22 years ago
|
||
Comment 71•22 years ago
|
||
Updated•22 years ago
|
Comment 72•22 years ago
|
||
Comment 73•22 years ago
|
||
Comment 74•22 years ago
|
||
Updated•22 years ago
|
Updated•22 years ago
|
Comment 75•22 years ago
|
||
Comment 76•22 years ago
|
||
Comment 77•22 years ago
|
||
Comment 78•22 years ago
|
||
Comment 79•22 years ago
|
||
Comment 80•22 years ago
|
||
Comment 81•22 years ago
|
||
Comment 82•22 years ago
|
||
Comment 83•22 years ago
|
||
Comment 84•22 years ago
|
||
Comment 85•22 years ago
|
||
Comment 86•22 years ago
|
||
Comment 87•22 years ago
|
||
Comment 88•22 years ago
|
||
Comment 89•22 years ago
|
||
Comment 90•22 years ago
|
||
Comment 91•22 years ago
|
||
Comment 92•22 years ago
|
||
Comment 93•22 years ago
|
||
Comment 94•22 years ago
|
||
Updated•22 years ago
|
Comment 95•22 years ago
|
||
Comment 96•22 years ago
|
||
Comment 97•22 years ago
|
||
Updated•22 years ago
|
Comment 98•22 years ago
|
||
Comment 99•22 years ago
|
||
Comment 100•22 years ago
|
||
Comment 101•22 years ago
|
||
Comment 102•22 years ago
|
||
Comment 103•22 years ago
|
||
Comment 104•21 years ago
|
||
Comment 105•21 years ago
|
||
Comment 106•21 years ago
|
||
Comment 107•21 years ago
|
||
Comment 108•21 years ago
|
||
Updated•21 years ago
|
Comment 109•21 years ago
|
||
Comment 110•21 years ago
|
||
Comment 111•21 years ago
|
||
Comment 112•21 years ago
|
||
Comment 113•21 years ago
|
||
Comment 115•21 years ago
|
||
Comment 116•21 years ago
|
||
Comment 117•21 years ago
|
||
Comment 118•21 years ago
|
||
Comment 119•21 years ago
|
||
Comment 120•21 years ago
|
||
Comment 121•21 years ago
|
||
Comment 122•21 years ago
|
||
Comment 123•21 years ago
|
||
Comment 124•20 years ago
|
||
Updated•20 years ago
|
Comment 125•20 years ago
|
||
Comment 126•20 years ago
|
||
Comment 127•20 years ago
|
||
Comment 128•20 years ago
|
||
Comment 129•20 years ago
|
||
Comment 130•20 years ago
|
||
Comment 131•20 years ago
|
||
Comment 132•20 years ago
|
||
Comment 133•20 years ago
|
||
Comment 134•20 years ago
|
||
Comment 135•19 years ago
|
||
Comment 136•19 years ago
|
||
Comment 137•19 years ago
|
||
Comment 138•19 years ago
|
||
Comment 139•19 years ago
|
||
Comment 140•19 years ago
|
||
Comment 141•19 years ago
|
||
Comment 142•19 years ago
|
||
Comment 143•19 years ago
|
||
Comment 144•18 years ago
|
||
Comment 145•18 years ago
|
||
Comment 146•18 years ago
|
||
Comment 147•18 years ago
|
||
Comment 148•18 years ago
|
||
Comment 149•18 years ago
|
||
Comment 150•18 years ago
|
||
Comment 151•18 years ago
|
||
Comment 152•18 years ago
|
||
Comment 153•18 years ago
|
||
Comment 154•18 years ago
|
||
Comment 155•18 years ago
|
||
Comment 156•17 years ago
|
||
Updated•17 years ago
|
Comment 158•16 years ago
|
||
Comment 159•16 years ago
|
||
Comment 160•15 years ago
|
||
Comment 161•15 years ago
|
||
Comment 163•11 years ago
|
||
Comment 164•10 years ago
|
||
Comment 165•10 years ago
|
||
Comment 166•10 years ago
|
||
Comment 167•10 years ago
|
||
Comment 168•10 years ago
|
||
Comment 169•8 years ago
|
||
Updated•8 years ago
|
| Comment hidden (advocacy) |
Comment 172•4 years ago
|
||
Just highlighting that this issue isn't dead. It is as important as ever to be able to load and save .mhtml files.
We need a better alternative to sharing documents than PDF files. MHTML files are going to be more accessible, and more convenient for people to share content.
See this appraoch to PDFs from the UK - https://gds.blog.gov.uk/2018/07/16/why-gov-uk-content-should-be-published-in-html-and-not-pdf/
Now all Chrome based browsers can load/save MHTML files to make it easy to share a snapshot of a site in time. Unfortunately FF users still need to install an extension..
Comment 173•4 years ago
|
||
(In reply to Mike Gifford from comment #172)
Now all Chrome based browsers can load/save MHTML files to make it easy to share a snapshot of a site in time. Unfortunately FF users still need to install an extension..
Can't someone just pull that code from the extension, and build it right into Firefox itself? Isn't this how a lot of programming happened here?
Comment 174•4 years ago
|
||
(In reply to Worcester12345 from comment #173)
(In reply to Mike Gifford from comment #172)
Now all Chrome based browsers can load/save MHTML files to make it easy to share a snapshot of a site in time. Unfortunately FF users still need to install an extension..
Can't someone just pull that code from the extension, and build it right into Firefox itself? Isn't this how a lot of programming happened here?
Or just pull the code from Chromium?
I'm sure such a feature would be welcomed by many users. :)
Comment 175•4 years ago
|
||
Mike Gifford <mike@openconcept.ca> wrote :
Now all Chrome based browsers can load/save MHTML files to make it easy to share a snapshot of a site in time. Unfortunately FF users still need to install an extension.
What? An extension? Did I missed something? I was under impression that since the notorious ver. 57 (web)extensions are no longer capable of adding a ability to view a format, that Mozillaʼs browser itself does not support; was I wrong?
Comment 176•4 years ago
|
||
For the purpose of saving complete webpage in one file, much more convenient file format would be HTML instead of MHTML, since HTML is widely-supported file type.
One of web extensions that already make this possible is called "Save Page WE". It has been actively maintained. Users can save a complete webpage or selected items. Further, users can save one or more selected webpages (i.e. selected browser tabs). Information bar at top of each saved HTML file is supported as well and can be enabled among settings. Web extension is available for Firefox - link and Chrome - link.
In this case it is questionable whether the development of native MHTML support makes sense nowadays – except for the purpose of opening old (archived) MTHML files. After all, this feature has been pending for more than 20 years.
Comment 177•4 years ago
|
||
I like Leon's pragmatic approach.
I think on the users scenarios, like dharing a page with somebody else.
One can send the link, but not always. For example when the page is local.
Also, the link is not useful when trying to share a form with some fields already loaded.
Sendint the HTML is a chore, it usually involves a number of files. Not useful for "normal" people. Let alone all the alarms raised by an email client receiving all those JS files.
People are sending each other MSWord or PDF files. The MSWord files don't have JS but VBA and fostered a terrible wave of destructive viruses (you opened the file and the Windows PC was nuked) because VBA had access to the OS bowels.
People are also sharing PDF files, which brings us back to the 1999 HTML pages that had fixed format and a gazillion spacer.gif refs.
For now, and 20 years later, people can't share web pages, or sets of web pages (what Leon mentions as a tabs set).
Ideally those complete web page(s) would travel encapsulated in a compressed file, optionally encrypted.
this would top the MSWord or PDF formats by being responsive and by including some of the link targets.
The issue is how to protext normal people from scams.
Like external links to bad places, or JS attacks.
I don't know, but I think that nowadays it should be easier than 20 years ago, when the VBA viroses were averriden.
Comment 178•4 years ago
|
||
Leon <leon.slo@outlook.com> wrote:
For the purpose of saving complete webpage in one file, much more convenient file format would be HTML instead of MHTML since HTML is widely-supported file type.
HTML is not a format for saving a complete page to a single file.
One of web extensions that already make this possible is called "Save Page WE". It has been actively maintained.
And another one (arguably, more actively maintained) is ‘Single File’.
In other words, by ‘format’ you mean specific hacks, employed by those extensions.
Widely-supported, you say? The width of their support among webbrowsers is a round sum of 0. Zero browsers support saving to that format. Please correct me, if I am mistaken.
In this case it is questionable whether the development of native MHTML support makes sense nowadays – except for the purpose of opening old (archived) MTHML files. After all, this feature has been pending for more than 20 years.
Please note, that if you count third-party addons, than prior Mozilla decided to trash all the labour of XUL-extension devs, they had managed to develop at least couple of addons, that implemented MHTML: one was ‘MAFF’, another had some self-descriptive name; so the issue was not that outstanding until year 2017.
In any case, I strongly doubt, that the lack of adoption of interchange format in a browser with usage share, that all recent years have been positively heading towards a statistical error, is a good measure to make judgements about its future.
The truth is that MHTML is simply the only saving format, supported by the most widely used browser on this planet: Google Chrome for Android.
Comment 179•4 years ago
|
||
(In reply to Dmitry Alexandrov from comment #178)
HTML is not a format for saving a complete page to a single file.
Not by original design, but it can be modified for this purpose with the help of extensions, as you mentioned.
Widely-supported, you say? The width of their support among webbrowsers is a round sum of 0. Zero browsers support saving to that format. Please correct me, if I am mistaken.
You are right, but I wasn't talking from the perspective of saving. My point was that you don't need an extension to be able to open single-file HTML. You only need such extension for saving a webpage into one file. In terms of universal ability for opening it, single-file HTML is recognized by every major web browser: tested with single-file HTML (saved by Save Page WE) on Firefox and Edge for Windows, as well as Chrome and Brave for Android.
Comment 180•4 years ago
|
||
Redirect a needinfo that is pending on an inactive user to the triage owner.
:hsinyi, since the bug has recent activity, could you have a look please?
For more information, please visit auto_nag documentation.
Comment 181•4 years ago
|
||
HTML is the most used language in the world.
Sadly, we all can read it (rendered) but comparatively few people can write and communicate it.
It should be mainstream-easy to write a page, or a linked pages cluster, using a tool similar to MSWord or the various WYSIWYG online editors, and then to send said page(s) wuth ease.
As the result of this lack of tools, people are sending each other PDF documents, an infamous printing format that generally requires to display a portrait A3-sized page in a landscape screen, and can't reflow its text(!).
IMO this MHTML format is the seed, a first step, to take over the documents communications world.
Updated•4 years ago
|
Updated•3 years ago
|
Comment 182•3 years ago
|
||
(In reply to Leon from comment #179)
(In reply to Dmitry Alexandrov from comment #178)
HTML is not a format for saving a complete page to a single file.
Not by original design, but it can be modified for this purpose with the help of extensions, as you mentioned.
The problem is that it is difficult for the average person to differentiate between single-page HTML and HTML pages saved as separate files. If they want to send an HTML file to a friend, they must either bundle up a bunch of files (which is beyond the technical expertise of many people), point them to a website, or just give up.
MHTML is a great format for sending web pages. Beside the obvious single-page format, IT IS SUPPORTED BY ALL MAJOR BROWSERS EXCEPT FIREFOX. (I don't count Safari, which also doesn't support it, because Macs don't count as computers.)
Google Chrome, Microsoft Edge, and Opera all open MHTML files natively, without extensions. WHY IS FIREFOX SO FAR BEHIND???
Widely-supported, you say? The width of their support among webbrowsers is a round sum of 0. Zero browsers support saving to that format. Please correct me, if I am mistaken.
You are right, but I wasn't talking from the perspective of saving. My point was that you don't need an extension to be able to open single-file HTML. You only need such extension for saving a webpage into one file. In terms of universal ability for opening it, single-file HTML is recognized by every major web browser: tested with single-file HTML (saved by Save Page WE) on Firefox and Edge for Windows, as well as Chrome and Brave for Android.
So...your argument is that people who are non-technical should have to learn how to install an extension in Firefox to do what EVERY OTHER MAJOR BROWSER CAN DO NATIVELY?!?
It seems much more reasonable to add support for saving HTML pages to a single-file directly into Firefox, in all major Web archive file formats, including MHTML/MHT, MAFF (old Firefox format), WARC (Internet Archive's Web ARChive format for entire sites), HTMLD (HTML Directory), and even Webarchive (in Safari browsers).
The default should be MHTML/MHT:
(1) It is widely supported.
(2) It doesn't end with HTML/HTM, so it won't be confused with those
(3) It is compatible with email messaging (being identical or almost identical to the EML format)
(4) It is a handy way to save web pages into a single-file and send them to others.
The only one that is better is the MAFF format, because it allows for compression. (I don't know if MHTML/MHT does or not, but it should...)
I was extremely angry when Mozilla removed the ability to save and read MAFF and MHTML/MHT formats. I've saved hundreds of pages in those formats and really don't want to change browsers or convert every single file to the new and less useful single-page HMTL format.
If security is a concern, then turn off this feature by default, but let users select it in Settings. Browsers should never become so secure that their usefulness disappears. Security must be coupled with ability for software to have value.
ADDING MHTML/MHT SUPPORT SHOULD BE A PRIORITY FOR FIREFOX!
Comment 183•3 years ago
|
||
(In reply to Leon from comment #176)
For the purpose of saving complete webpage in one file, much more convenient file format would be HTML instead of MHTML, since HTML is widely-supported file type.
The problem with your argument is that complete page HTML is not distinguished from HTML which is scattered into dozens and dozens of images and folders and files. That makes it confusing for the average user who just wants to send a webpage to someone else. Links are inadequate and single-file HTML works, but to repeat the obvious, it gets confused with non-single-file HTML.
So the single-file HTML is NOT useful in this case. It's the opposite of useful.
One of web extensions that already make this possible is called "Save Page WE". It has been actively maintained. Users can save a complete webpage or selected items. Further, users can save one or more selected webpages (i.e. selected browser tabs). Information bar at top of each saved HTML file is supported as well and can be enabled among settings. Web extension is available for Firefox - link and Chrome - link.
Save Page WE is a fantastic extension! However, it still doesn't address the main problem... Single files to send to other people should be named with something OTHER than HTML/HTM. Otherwise, normal users will not send those files correctly, losing images and other information from a web page.
In this case it is questionable whether the development of native MHTML support makes sense nowadays – except for the purpose of opening old (archived) MTHML files. After all, this feature has been pending for more than 20 years.
MHTML/MHT makes fantastic sense today!
At this point, we can save web pages as "Web Page, complete" or "Web Page, HTML only" natively in Firefox and the HTML single-file format from the "Save Page WE" or "Singlefile" extensions.
"Web Page, complete" saves dozens and dozens of objects...for this page it saved 162 separate files and/or folders.
"Web Page, HTML only" doesn't save anything but the basic HTML, losing all the images and other important features on a web page.
The single-file HTML format from the extensions is great...but it saves with the HTML extension. So that means that people who want to send a file or open the file can get easily confused. It also doesn't allow for compression, which would be tremendously useful.
Thus it is proved that MHTML is very useful for many, many people, since normal users who could use this feature represents the vast majority of browser users.
Plus, I really like it too.
Comment 184•1 year ago
|
||
(In reply to Biju from comment #83)
And thunderbird can save *.eml file, and also can view *.mht and *.eml files
This comment is 21 years old and still valid, as far as I understand.
So TB has this functions out of box but FF can't view mht and needs third-party addon for saving.
And there is no more common option to save a page to single file (pdf is not capable of complex layout).
Comment 187•1 year ago
|
||
I recently rediscovered the MHTML single-file format while looking for a better way to save AI conversations, and it turns out to be incredibly useful for my needs. I often have interesting discussions with AI, so much so that I’ve resorted to saving some of them in .docx files. However, the formatting was far from ideal.
Upon discovering MHTML, I realized it perfectly suits my use case: it’s fast and convenient (Ctrl+S, name the file, done) compared to opening a word processor and manually copying and pasting content. More importantly, the final result closely preserves the original appearance of the live webpage, making it much more readable.
Given how long this format has existed, I was genuinely surprised that it hasn’t been implemented yet. I strongly support its adoption in Firefox.
Comment 189•1 year ago
|
||
This is the format we all should be using, instead of the ancient PDF that forces us to read portrait pages in landscape screens without a chance to reflow the text.
Not only for @vlakoff's use case, but for everything except the most boring legal docs, those we all skip.
I think that we could start by binding the MHTML format to the PC's default browser, so those files could be opened with a double click.
And perhaps having a chrome less browser choice; in this case less might we'll be more.
Comment 190•1 year ago
|
||
Check comment https://bugzilla.mozilla.org/show_bug.cgi?id=40873#c165
about the same as above, 10 years ago
| Assignee | ||
Comment 191•8 months ago
|
||
Hey Folks (Emilio, Smaug and Baku)!
I'm interested if Mozilla would accept a patch here if someone worked on it.
- I've got a working patch that I hacked together to exporting .mhtml files, it adds an export file type in the save menu.
- I think it's worthwhile just landing export if we're happy with it and I'd be happy to work in it.
- The patch I've worked on needs some more validation, I've tested only a couple of pages but wanted to validate if it'd be accepted first. I'm also not expecting this patch to be close to what should be landed either, happy to discuss further direction.
On my travels, I spotted that there's a bunch of unanswered issues:
- Handling of @ import doesn't always work in Chromium
- Shadow DOM changes seems a little hacky.
I think they're ok, to leave as is. This isn't a regular web standard that will impact a significant number of users. I think issues could be resolved iteratively.
The follow up work here would be:
- Supporting mhtml viewing, because of prior Chromium security issues in viewing I think a different origin attribute might be worth considering to reduce storage exfiltration.
- Adding the extension API to download mhtml snapshots from the background script.
Thanks!
Comment 192•8 months ago
|
||
(In reply to Jonathan Kingston [:jkt] he/him from comment #191)
Thanks for looking into this Jonathan!
I'm interested if Mozilla would accept a patch here if someone worked on it.
I don't have any problem with it in principle, but I don't have the whole context here. Looking a bit at the backscroll here I don't see any obvious objection, so seems to me worth experimenting with / landing behind a pref at least. We probably want to check with product before enabling by default but that can be done alongside the implementation, and as long as the implementation isn't super-invasive I don't expect others to object...
- Handling of @ import doesn't always work in Chromium
- Shadow DOM changes seems a little hacky.
Yeah serialization of shadow dom is always fun. But it seems we can use declarative shadow dom nowadays for that if needed?
I think they're ok, to leave as is. This isn't a regular web standard that will impact a significant number of users. I think issues could be resolved iteratively.
I agree they're probably not blockers for an initial implementation.
The follow up work here would be:
- Supporting mhtml viewing, because of prior Chromium security issues in viewing I think a different origin attribute might be worth considering to reduce storage exfiltration.
Should this really be a follow-up? It seems to me it might be worth doing it the other way around (first implement viewing, then exporting)? Exporting something we can't see isn't quite useless, but... In any case landing some initial version of the export code seems fine since you already got started on it, but I would expect turning it on by default to be blocked on us being able to display the file...
Regarding the origin stuff, should we just use an opaque origin (like file:// or data: uris) instead? Not sure we need a new origin attribute, you conceptually want each individual mhtml archive to be "isolated" in a way, right? An opaque origin gives you that.
| Assignee | ||
Comment 193•8 months ago
|
||
Should this really be a follow-up? It seems to me it might be worth doing it the other way around
Oh I agree on the framing of blocking release for sure. Certainly a pref to start with too. I agree reading would make more sense to start with, I'm most interested in exporting crawls from extension context if I'm honest (kinda was sucky to write an extension that's Chrome only and this is a blocker). Essentially I wanted it as an archival format, I'm not really reading it at all (in fact I'm converting it back to html for my use-case too).
It seems like Chrome uses file extension mapping to mthml parsing after mime checking it additionally preventing all scripts and only allowing in file:// loads.
Not sure we need a new origin attribute, you conceptually want each individual mhtml archive to be "isolated" in a way, right?
Yeah fair, I'll check their prior security issues some more too. I think restricting script loads or it's ability to load network requests at all is closer to what model that's suitable.
Yeah serialization of shadow dom is always fun. But it seems we can use declarative shadow dom nowadays for that if needed?
Good point yeah! I think we can iterate on this rather than expect it all to be supported in one patch. (which it seems like you're happy with too).
I think there's enough test cases we can switch to WPT to get a working solution.
In terms of alternatives here there's WARC/WACZ that wget and http archive use, it's much heavier on implementation (it's sort of similar to MHTML+HAR in fidelity) the implementation would be radically different and similar to HAR would need to be a dev tool only opt-in due to performance and complexity.
Anyway, thanks for your attention here. I'll continue with implementing a prototype that does import and export.
Comment 194•8 months ago
|
||
I think there should be at least some ideas how to implement viewing before adding exporting, since I don't think we should add only the latter.
Then adding exporting behind a pref would be fine.
dmcintosh has been looking into this too recently.
| Assignee | ||
Comment 195•8 months ago
|
||
| Assignee | ||
Comment 196•8 months ago
|
||
Broadly the viewing support is:
-
MHTMLStartup.sys.mjs
Registers .mhtml/.mht → "multipart/related" via category manager -
components.conf
Registers MHTMLDocumentLoader as Gecko-Content-Viewer for "multipart/related" -
When Firefox loads an MHTML file:
docshell → content type = "multipart/related" → MHTMLDocumentLoader.createInstance() -
MHTMLDocumentLoader:
- Changes content type to "text/html"
- Gets the standard HTML document loader factory
- Wraps it with MHTMLStreamListener
-
MHTMLStreamListener:
- Buffers all MHTML data in onDataAvailable()
- In onStopRequest(): parses MHTML, converts to HTML with embedded data: URLs
- Feeds converted HTML to the underlying HTML viewer
I did have a DocShell implementation but this seems much cleaner and is working.
Comment 197•8 months ago
|
||
Subresource loading (including iframes) and principal handling are perhaps the trickiest things to figure out.
Comment 198•8 months ago
|
||
I was kind of assuming that they would be treated virtually like data uris...
Comment 199•8 months ago
|
||
But data urls aren't virtual, they are there in the source code, in src attributes or so. mhtml is different, no? It defines some resources in the .mhtml.
| Assignee | ||
Comment 200•8 months ago
|
||
But data urls aren't virtual, they are there in the source code, in src attributes or so. mhtml is different, no? It defines some resources in the .mhtml.
They support cid:, relative and absolute paths. The latter doesn't seem super common but I think it's handled by Chrome. cid: are references to the embedded mime chunks.
cid: I think can just be a new protocol handler that knows where to get it's data from.
The absolute/relative HTTP paths, I was thinking would probably be best handled with a custom principal or load info that routes it to the document also.
Comment 201•8 months ago
|
||
This what I get for not CC'ing myself on this! :P Thanks smaug for the needinfo, hadn't realized things were happening here. I'll comment quickly before I can take more of a look later.
I've been working on opening MHTML files a bit over the last few weeks as an exploratory/learning thing, only the opening part though. I've uploaded what I've made so far to GitHub; it seems like my approach is pretty similar to what you've done. I haven't read much of your patch yet, I'll try to take a look later today.
Going off of your comment, it looks like we both take the same approach of intercepting the content (I used nsIStreamConverter instead of Gecko-Content-Viewer, not sure how they differ?), then translating URLs in the document into data URIs. I don't like this approach (that's why I reached out to smaug), but it does kind of work. I feel like there's something with synthesizing responses like service workers do that might work better, or maybe your idea with custom principals or loadinfo might work too, I'm not familiar with any of this. (I usually work on frontend-y things and the installer :) )
My parser is also a lot longer (and a bit ugly internally), but I do like how it can issue warnings for some issues with the file—Chrome's 'malformed multipart archive' annoyed me. Might be worth mixing the two patches and tests.
Re shipping this, I've been talking to product a little bit about it—from what I can tell the thinking is that shipping MHTML for the sake of shipping MHTML might not fly, but using it to satisfy some usecases might work. I'll keep talking about that internally, especially if these patches start coming together. I'll note that the opening half is technically bug 18764, which was wontfix'd due to not being prioritized, but internally I haven't really seen much opposition to the concept.
| Assignee | ||
Comment 202•8 months ago
|
||
Going off of your comment, it looks like we both take the same approach of intercepting the content (I used nsIStreamConverter instead of Gecko-Content-Viewer, not sure how they differ?),
Yeah I think the DocShell approach I had was much more mature in that sense, it intercepted the requests properly and assigned an opaque origin to the loads. I gave up because it had trouble loading the files and felt a little heavyweight but I think the current approach is far too simplistic.
Might be worth mixing the two patches and tests.
Yeah not protective, I think having ref rendering tests between mthml and html would be nice.
from what I can tell the thinking is that shipping MHTML for the sake of shipping MHTML might not fly,
I mean it's only going to be visible as a single export option in save. My expectation is it could be argued as an extension parity case. I think if the implementation isn't too cumbersome then that's enough justification in my mind.
I feel like there's something with synthesizing responses like service workers do that might work better, or maybe your idea with custom principals or loadinfo might work too, I'm not familiar with any of this. (I usually work on frontend-y things and the installer :) )
I think I can continue to explore these and we can compare and contrast easily.
but using it to satisfy some usecases might work.
My use case is refine.page a currently Chrome only extension that there's kinda no alternatives, Single File is the only close alternative, there's loads of reasons not to use it imo.
Additionally I wanted this only a few months back for RAG'ing page content, the snapshots of pages via an extension is a good stable format to use for that.
| Assignee | ||
Comment 203•8 months ago
|
||
Yeah modelling it on loadinfo seems the cleanest for a more indepth approach that isn't doing parsing.
I was thinking the ID could be stored on the principal but they're not really designed for storage and more load info. I can then use that to store an index into a global registry that keeps track of mhtml loads.
I feel like there's something with synthesizing responses like service workers do that might work better, or maybe your idea with custom principals or loadinfo might work too
The current approach I'm using seems in this direction. It's a bit like a blob handler for the document itself which then intercepts the document requests like a worker does.
| Assignee | ||
Comment 204•8 months ago
|
||
| Assignee | ||
Comment 205•8 months ago
•
|
||
I need to do a bunch more validation (namely I didn't check iframes, I think we might need to propagate the loadinfo flags downwards into the child frame), testing and reading but I'm pretty happy with this approach now.
I explained this new approach more above. It'd be good to validate with product regarding what would block landing this type of work. We could split out the export, but we'd lose some ability to do comparison tests of HTML->MHTML and then reftesting those two files.
I think the only real missing part now is adding a pref to wrap it.
Comment 206•8 months ago
|
||
Wow, thanks!
I've taken a quick look at the opening half of the patch, and it looks pretty neat. Main question: what commit is it based on? :) I tried applying it, and it didn't seem to apply to 43c97a06 or 33bba5cf with git apply, but I'm also not used to applying patches off Bugzilla so maybe I'm doing it wrong. I'll look at it more once I can apply it properly.
I tried implementing the opening part of your approach based on your patch, and I'm a little unsure how it works—I think? nsDocShell::ShouldIntercept et al are only called in the parent process, but it looks like the actual MHTMLArchive storing the contents is only in the child process, which'd mean nothing gets intercepted...? My attempt doesn't seem to actually intercept requests, but my approach is quite different so it could work out in yours.
I'm also a little unsure whether nsDocShell's implementation of nsINetworkInterceptionController is the right place to put this, since e.g. FetchDriver also implements that interface, but I figure that's a 'figure that out once it works' problem.
Lastly, I'm thinking that we might want to avoid naming the classes 'MHTML' unless they literally deal with MHTML syntax, since it's possible other uses are found for the code and/or other formats are supported. If you have suggestions for a better name, let me know; I've been going with OfflineArchive, but I don't know if that's much of an improvement. Certainly can wait until a better name arrives. (SavedArchive? RestrictedViewOfTheWeb?)
Re product, I'll be speaking to them tomorrow so I can discuss it then.
| Assignee | ||
Comment 207•8 months ago
|
||
Sorry haven't had chance to revalidate the patches some more.
what commit is it based on
b82cded8c5b732c2ea15b7871d14e13b5fadeffd it should be based on. I can make you a git commit instead, might be me as it's been a while since using hg so I'm using git as a the flow we used to use, forgive my external contributor noise if it's me.
nsDocShell::ShouldIntercept et al are only called in the parent process, but it looks like the actual MHTMLArchive storing the contents is only in the child process, which'd mean nothing gets intercepted...
I think you're right, sorry. The test case I had was all embedded data/file URLs. Let's get a test case written to validate sub resource loaded in some complex cases.
My intent here was to have this handled in the parent process and then IPC message though to child process.
I'm also a little unsure whether nsDocShell's implementation of nsINetworkInterceptionController is the right place to put this, since e.g. FetchDriver also implements that interface, but I figure that's a 'figure that out once it works' problem.
Yeah let's get this working in this format, then shuffle things about.
Lastly, I'm thinking that we might want to avoid naming the classes 'MHTML' unless they literally deal with MHTML syntax,
Yeah I was thinking this too. Apple uses WebArchive, there's some room for overloading that or riffing from it perhaps. Chrome's extension API here is chrome.pageCapture.saveAsMHTML
Comment 208•8 months ago
|
||
(In reply to Jonathan Kingston [:jkt] he/him from comment #202)
My use case is refine.page a currently Chrome only extension that there's kinda no alternatives, Single File is the only close alternative, there's loads of reasons not to use it imo.
Additionally I wanted this only a few months back for RAG'ing page content, the snapshots of pages via an extension is a good stable format to use for that.
I'm very happy to see a native implementation which is better.
For now, I use SingleFile to save a lot of pages, it works very well.
But what are the reasons not to use it?
| Assignee | ||
Comment 209•8 months ago
|
||
But what are the reasons not to use it?
- License if you're bundling it.
- It's not quite as stable as using the browser to do this work. I had a bunch of edge cases in a directory, I couldn't find fixes for.
Comment 210•8 months ago
|
||
(In reply to Jonathan Kingston [:jkt] he/him from comment #209)
But what are the reasons not to use it?
- License if you're bundling it.
- It's not quite as stable as using the browser to do this work. I had a bunch of edge cases in a directory, I couldn't find fixes for.
Thanks for the answers, I feared you wanted to say that they track my visited pages and collect to much data.
I really hope your code gets merged and it's much faster than SinglePage.
| Assignee | ||
Comment 211•8 months ago
|
||
Backup of moving the patch to the parent process. This seems to be working much better now.
| Assignee | ||
Comment 212•8 months ago
|
||
This fixes rendering of CID, I've not seen any disparity with Chrome now.
| Assignee | ||
Comment 213•8 months ago
|
||
Duncan does it make sense to sync on this, I think we're close to being landable. What does product say?
Additionally in:
- https://issuetracker.google.com/issues/41344244 I have a patch to serialize shadow DOM into mhtml, this works really quite well over the existing Chromium approach. I think we should add that approach to the Firefox patch from day one tbh (assuming it doesn't bloat it considerably).
- Somewhat unrelated I started an explainer on making <canvas> serialize better too: https://github.com/WICG/proposals/issues/258
| Assignee | ||
Comment 214•8 months ago
|
||
Comment 215•8 months ago
|
||
Thanks for working on this and submitting the Phabricator revision!
I feel like the serialization improvements (the CSS url extraction and your shadowdom/canvas work) might be good, but they might be better done in parallel under a different bug so they don't block MHTML and vice versa, since I suspect the patches don't overlap much.
I tried building/testing your patch. Saving seems to work at least at a first glance; images on www.wikipedia.org didn't seem to get saved but that's a relatively minor issue given that the rest of the page saved well. The main issue is that I can't seem to open MHTML files? The tests in toolkit/components/windowcreator/test can't seem to either, in both cases it just opens the 'what do you want to do with this file' dialog. How are you running Firefox/do you have other patches? I was also thinking that nsDocShell's nsINetworkInterceptController only runs for docshells in the parent, which isn't most of them, so it shouldn't intercept much... do you have e10s off for some reason?
I've been working on my own branch for viewing specifically, you can see the full diff on GitHub if you're interested. The resources are made available by a property on the document, which is exposed by native IPC to the parent process for interception, where the intercept controller is attached to the ParentChannelListener. It still needs work, but the initial feedback that I've got seems positive, and it worked well for a demo I did on Tuesday. I'm hoping to refactor MHTMLParser a bit shortly, add documentation and iframe support, etc.
My current inclination is to keep your saving code (since the patch already works and I suspect it's the more important part to you) and use my opening code (since I can get it to work, I think the design is a bit better, and it's more important to me :) ), but I'm open to something else. I'll probably submit a more complete stack of patches in a few days to bug 18764, and I'll leave some comments on your patch for a few things I noticed.
Again, thanks for working on this—even though my code is pretty different, there's no chance I'd have gotten there without your patches.
Re product—they had incomplete information on the security properties, so if everything's good there then there might not be much problem. User needs were (a) to deal with pushback internally if that appears, and (b) because MHTML will realistically be a power-user-y feature, which is OK but there might be interesting things that can be done with the underlying platform feature, in my patch the nsIOfflineArchive stuff and in yours the MHTMLArchive stuff. My overall understanding is that a solid implementation that passes some sort of security scrutiny should be ok to ship.
(I'll also note that the 'hack cycle' during which I was working on this is over, but hopefully I can still find time to push this forward.)
| Assignee | ||
Comment 216•8 months ago
|
||
| Assignee | ||
Comment 217•8 months ago
|
||
| Assignee | ||
Comment 218•8 months ago
|
||
I feel like the serialization improvements (the CSS url extraction and your shadowdom/canvas work) might be good, but they might be better done in parallel under a different bug so they don't block MHTML and vice versa, since I suspect the patches don't overlap much.
Yeah lower priority for sure. Especially the canvas thing that's a whole discussion. Would be great to get your input.
opens the 'what do you want to do with this file' dialog.
I think that issue sounds like the the MIME type mapping or document loader. Let me validate the moz-phab has everything on my local. I've just updated it to latest central, clobbered and still works fine.
in the parent, which isn't most of them, so it shouldn't intercept much... do you have e10s off for some reason?
Nah it's on. I've been able to import and export the files from Chromium and back.
Broadly my read approach is:
- MHTMLStartup registers .mhtml/.mht extensions → multipart/related MIME type via category manager.
- MHTMLDocumentLoader is registered as a Gecko-Content-Viewers for multipart/related.
- When Firefox opens an MHTML file, the document loader factory intercepts it, parses the MHTML, and feeds the HTML directly to the standard HTML document loader.
I just rebuilt and was able to see Wikipedia images into Chromium (and Chromium ones in my build).
Are you doing debug builds, I'm not. I'm running on MacOS. Attached some screenshots.
but I'm open to something else
Not super fussed, as you say export was more interesting to me. I can follow up with the extension API once landed. Maybe ask someone else the best direction to take. I suspect there's tests from both of us that are worth taking. Additionally we might want to move them to wpt tests too.
Re product... might be interesting things that can be done with the underlying platform feature
Great! Yeah, the most interesting thing I see is resolving the export via extensions which as I say I've wrote two recently.
Comment 219•8 months ago
|
||
(In reply to Jonathan Kingston [:jkt] he/him from comment #218)
in the parent, which isn't most of them, so it shouldn't intercept much... do you have e10s off for some reason?
Nah it's on. I've been able to import and export the files from Chromium and back.
I just built your patch off of Phabricator (onto the main branch, 28398d11 which is a bit after your patch), and I still can't open MHTML files. I still think there's some multiprocess thing going on—I noticed in your tests, e.g. browser_mhtml_load.js:
let doc = tab.linkedBrowser.contentDocument;
ok(doc, "Document should be loaded");
but contentDocument only works if the browser is in-process, which it shouldn't be for a file/http page. (I ran into this problem in my tests and had to do a bunch of SpecialPowers fun.) If you open the browser toolbox and type gBrowser.selectedBrowser.contentDocument, what do you get? I'd expect null except on about: pages.
I just rebuilt and was able to see Wikipedia images into Chromium (and Chromium ones in my build).
Looks like images do work on Wikipedia articles, but not ones around the UI. I meant https://www.wikipedia.org/ proper, and I see in your screenshot that the hamburger menu is missing from the Firefox-saved one. A glance at the inspector suggests they use url() for some reason, so that part of your patch might not be working right? Probably not an MHTML problem though, and doesn't change that the rest of the page saves fine :).
Are you doing debug builds, I'm not. I'm running on MacOS. Attached some screenshots.
I'm on Windows with non-debug builds, I wouldn't expect the debug builds to break it though?
I also meant to leave some comments on your Phabricator revision, I'll go do that now.
| Assignee | ||
Comment 220•7 months ago
|
||
but contentDocument only works if the browser is in-process, which it shouldn't be for a file/http page. (I ran into this problem in my tests and had to do a bunch of SpecialPowers fun.) If you open the browser toolbox and type gBrowser.selectedBrowser.contentDocument, what do you get? I'd expect null except on about: pages.
Null yeah. I'll investigate what the test is doing. I think that's why they're marked as skipped currently, they predate the last rewrite.
Looks like images do work on Wikipedia articles, but not ones around the UI. I meant https://www.wikipedia.org/ proper, and I see in your screenshot that the hamburger menu is missing from the Firefox-saved one. A glance at the inspector suggests they use url() for some reason, so that part of your patch might not be working right?
I think this is the parser parts that in phab you asked if could be split out. I'm wondering if there's a better approach overall there. Removing that code certainly breaks more things.
I'm on Windows with non-debug builds, I wouldn't expect the debug builds to break it though?
I wouldn't no. Just checking an assert wasn't giving the different behaviour.
Thanks for the continued reviews. I'll upload the commit split I just made but there's some issues I need to investigate more as mentioned.
Updated•7 months ago
|
| Assignee | ||
Comment 221•7 months ago
|
||
Adds the ability to save web pages as MHTML (single-file archive).
Components added:
- nsMHTMLPersist: Core MHTML generation with MIME multipart encoding
- nsWebBrowserPersist changes: Integration with persist framework
- MHTMLArchive.sys.mjs: JS-level archive creation
- contentAreaUtils.js: Save dialog MHTML option
- PERSIST_FLAGS_SAVE_AS_MHTML flag for enabling MHTML output
Includes roundtrip tests to verify reading and writing work together.
Depends on D279105
| Assignee | ||
Comment 222•7 months ago
|
||
Currently working on stripping out any parsing logic at all, other than the mime parsing. I think all of that is just hacks working around a problem that should be a little simpler.
I've pushed it to gh, it'll need refactoring back into the split commits: https://github.com/mozilla-firefox/firefox/compare/main...jonathanKingston:firefox:jkt/40873-remove-parse?expand=1
To avoid any stripping or changing of the parser I've added a media query to hide things like <noscript>
Once finished I'll upload it again to moz-phab.
| Assignee | ||
Comment 223•7 months ago
|
||
I pushed the updates to moz-phab.
I've have removed any parser parts fully, I've also used an export change to drop <noscript> elements from being exported to fix exporting to Chrome.
Comment 224•7 months ago
|
||
On the saving side, I've left some comments on your Phabricator revision. I'm not sure why you seem to remap all of the resources to their CIDs, a few more tests of the saving might be helpful, and there's some cleanups and things to move around, but once those are addressed if you think it's ready for review you can probably just request review. (I'm guessing you have an idea of who good reviewers would be, and I'll add others if needed.)
I might suggest changing the revisions so they don't depend on each other, since in any case the viewing side will likely need more work. You mention having roundtrip tests, which is a good idea, but maybe they should go into a follow-up revision so saving doesn't have to wait for viewing?
On the viewing front, I talked to :asuth and he suggested against using nsINetworkInterceptController since he sees that as a service-worker-specific thing that outside code shouldn't use, despite the comment in that IDL suggesting it's more general. I've ported my viewing code over to using redirects to the CID protocol, which is now working!, but I haven't cleared that by Necko developers yet. I also was figuring that I think it could be intercepted in the IOService, in which case we could match Chromium's behaviour on non-HTTP URLs and avoid cid: needing to be valid outside of that context, but I haven't explored that one yet; I plan to ask some Necko developers on Monday as to what the best approach would be, and I'll update this bug when I find out.
(In reply to Jonathan Kingston [:jkt] he/him from comment #223)
I've have removed any parser parts fully, I've also used an export change to drop <noscript> elements from being exported to fix exporting to Chrome.
I like the new <noscript> handling much better than before, thanks for working on it! I'd like to clarify some of your patches' behaviour around the tags, though: my understanding as for why to remove them when saving is that we want the page to display as it was rendered, which means not displaying noscript if they wouldn't have been displayed at the time. In this case, I agree with removing them, although we maybe should only remove them if scripting is enabled for the document.
On the viewing side, though, I don't agree with the CSS/parser change that hides them if they're already in the file. If Chromium already removes them, which it seems to, then in my eyes most MHTML files already don't have these tags, and if the tags are there then—since scripting is disabled, and it is just rendering the HTML page—it should render them as usual. This also would match Chromium's behaviour, where noscript tags are rendered if present. Is there a specific reason to hide them?
In any case, thanks for your patience and continuing to work on this!
Comment 225•7 months ago
|
||
I've gotten word that modifying NS_NewChannel/nsIOService can work, and I've updated my GitHub branch to match, although it currently doesn't do stylesheet load progress correctly for some reason. (FYI I'll be out next week.)
| Assignee | ||
Comment 226•7 months ago
|
||
Sorry have been meaning to split the patches into separate bugs, then move away from nsINetworkInterceptController. I probably will have time tomorrow.
Comment 227•5 months ago
|
||
Has this bug split happened? Not sure where to look for this.
| Assignee | ||
Comment 228•5 months ago
|
||
I've created both: Add MHTML (RFC 2557) export support to "Save Page As" and Add MHTML (RFC 2557) read support for .mhtml / .mht files to separate out from here. I'll rebase and move my patches there.
We likely will continue to want follow up work here for example there's:
- Testing that validates export works with import to produce the same HTML rendered layout.
- Shadow dom changes that I'm aiming to also get landed in Chromium: https://issuetracker.google.com/issues/41344244
- Implement chrome.pageCapture for extensions.
Description
•