Open Bug 2055405 Opened 2 months ago Updated 7 days ago

Support MathML in tagged PDF generation

Categories

(Core :: Disability Access APIs, enhancement)

enhancement

Tracking

()

ASSIGNED

People

(Reporter: Jamie, Assigned: Jamie)

References

(Blocks 1 open bug)

Details

Attachments

(9 files, 1 obsolete file)

90.20 KB, application/pdf
Details
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review
48 bytes, text/x-phabricator-request
Details | Review

We generate tagged PDF for HTML now. It'd be nice if we could support MathML as well.

There are several ways to tag MathML in PDF. My initial thinking is that the simplest for us will be to include the MathML tags into the structure tree (SE method), rather than using the associated file (AF) method. There are a few things to solve to achieve that:

  1. We need to specify the MathML namespace on MathML tags. However, SkPDF doesn't support specifying the namespace for structure elements.
  2. We need to wrap the <math> element inside a PDF Formula structure element, even though there's no equivalent in the Gecko accessibility tree.
  3. We need to map the math roles in the Gecko accessibility tree to the correct MathML tag names.
  4. Eventually, we need to support attributes as well, but Gecko doesn't actually expose them in the parent process cache right now, so that will require additional work.

I have a WIP branch which does all of this. It correctly tags the MathML and pdf.js can read it. Unfortunately, a snag I've hit is that for some reason, the content inside the MathML tree can't be seen as text in the PDF. For example:

data:text/html,<p>the answer is <math><mfrac><msqrt><mi>x</mi></msqrt><mrow><mn>2</mn><mi>i</mi></mrow></mfrac></math> okay?

I don't see x, 2 and i anywhere when I load the exported PDF with pdf.js, even though I can see them in the layout frame tree when I load the markup. We're losing the text somehow when it gets painted and rendered.

If I can't solve that, I guess we'll have to use the associated file method instead. However, that's going to be a lot harder to support in SkPDF.

This seems to be dependent on the font that's being used for those math characters. When we're using STIX Two Math, they show up visually in the PDF output but they have a /ToUnicode mapping that maps them to U+0000, and so they don't appear properly in text copied from the PDF (and presumably things like search would similarly fail). But if I remove STIX from my Linux machine, they fall back to DejaVu Serif, and then they copy/paste out of the PDF as the expected characters.

This looks like an issue with this font (and presumably may affect others...) and the SkPDF backend, independent of the tagging issue.

Note that with SkPDF disabled, so that we use the cairo backend instead, I see a similar-but-not-identical issue: trying to copy the text from the generated PDF gives me U+FFFD REPLACEMENT CHARACTERs for the math letters & digit. So cairo-pdf also struggles with this font.

Also interesting to note that the radical character U+221A survives fine; it's just the letters and digits that are failing, though they're all coming from the same underlying STIX Two Math font.

:jfkthame, I tried on Windows and I've the same issue with CambriaMath but only when the glyphs are used as super/sub-scripts.
I don't know how SkPDF builds the ToUnicode map but maybe it's just wrong when the glyphs are from ssty and then it should fallback on the glyph which has induced the substitution.
Anyway I tried with Chrome and it works well.

:Jamie, in pdf.js we support either SE or AF. Out of curiosity why is it more complex to implement AF ?

:David, out of curiosity, in general, is there a better approach between SE and AF ?

Flags: needinfo?(d.p.carlisle)

:David, out of curiosity, in general, is there a better approach between SE and AF ?

most systems find it (a lot) easier to generate AF as they can take an existing generator from Word or tex or whatever to MathML and just attach generated MathML as a single blob to the /Formula element. Whereas to generate SE you need to essentially interleave the MathML through the PDF layout. However in your case you are in a somewhat better position as you are (I assume) starting from MathML input.

In theory MathML SE gives the potential for synchronizing the reading with the visual system cursor when navigating subterms because the SE is tied to each term in the formula, whereas AF is just associated with the formula as a whole, but as far as I know that isn't supported yet in pdf readers.

The tricky parts of the SE approach are arranging to artifact all the content that doesn't relate to the markup. For AF you put the MathML on the Formula so whatever content is used to render the math can be ignored by AT, But with SE if you have for example a square root, that's <msqrt><mi>x</mi></msqrt>which already produces the square root sign and an over-rule, so you need to mark them as Artifact so they don't end up in the structure tree. Similarly large stretched brackets around matrices you need to artifact the vertical glyph constructs so they don't end up in the structure tree while supplying "normal"(to put in the<mo>` elements in the MathML. More on that in the recently produced best practice guide for tagging math

https://pdfa.org/resource/best-practice-guide-math-in-pdf/

er I seem to have written more than I intended, I hope I understood the question and it's somewhat related to what you intended to ask:-)

(In reply to Calixte Denizet (:calixte) from comment #3)

:Jamie, in pdf.js we support either SE or AF. Out of curiosity why is it more complex to implement AF ?

SkPDF supports building a structure tree, so we effectively translate our accessibility tree into the right data structure and pass it onto SkPDF. Since MathML is just another subtree in the accessibility tree, we just need to translate the MathML nodes to the right MathML tags. The only thing we need to add to SkPDF is the ability to specify the namespace for a structure element.

In contrast, AF is an entirely new concept that SkPDF has no support for at all. So, exposing MathML via AF first requires teaching SkPDF about AF and then figuring out some way of plumbing that data through the SkPDF API. I'm not saying it's completely infeasible; it's just much harder and I don't have any clue how I'd do it just yet.

(In reply to David Carlisle from comment #4)

In theory MathML SE gives the potential for synchronizing the reading with the visual system cursor when navigating subterms because the SE is tied to each term in the formula, whereas AF is just associated with the formula as a whole, but as far as I know that isn't supported yet in pdf readers.

Even though synchronised reading isn't supported in readers, do you know if SE is as well supported as AF, or do some readers only support AF? If SE isn't widely supported, we may be better off putting the time into implementing AF anyway to benefit the maximum number of readers.

The tricky parts of the SE approach are arranging to artifact all the content that doesn't relate to the markup. ... with SE if you have for example a square root, that's <msqrt><mi>x</mi></msqrt>`
which already produces the square root sign and an over-rule, so you need to mark them as Artifact so they don't end up in the structure tree.

Yeah, I'm already seeing this problem (the square root sign appearing in the structure tree) and I haven't figured out how to deal with it just yet. I imagine the square root sign is painted somewhere specific, so I should in theory be able to mark it as an artifact, but I need to work out where that actually happens.

(In reply to James Teh [:Jamie] from comment #5)

Yeah, I'm already seeing this problem (the square root sign appearing in the structure tree) and I haven't figured out how to deal with it just yet.

I think I found a solution for this. It's in my WIP.

Calixte, I'm seeing aria-owns in the math subtree generated by my WIP with something like:
data:text/html,<math><msqrt><mi>x
This breaks Windows screen readers, which can't handle aria-owns in math subtrees as I explained in bug 1998046 comment 9. I thought bug 1998046 was fixed... but was it perhaps only fixed for the AF method? Would fixing this for SE be infeasible?

Flags: needinfo?(cdenizet)
See Also: → 1810914

I might be barking up the wrong tree here, but this seems to be similar to bug 1810914 in that we ultimately don't get the ideal ToUnicode mapping. As bug 1810914 notes, we could solve this using ActualText, but the problem is that we've lost the actual text by the time we get to drawing. There might be some less intense way to fix the mapping in this case though.

Out of curiosity, I took a run at getting Claude to plumb the actual text through to Skia. It did work, but it's not at all a small change and I am not familiar enough with any of this to be confident that it's the correct (or sufficiently performant) solution.

An alternative could be to specify the ActualText in the struct element for a text leaf node in a MathML subtree and have SkPDF pick that up when it's writing the marked content for the text leaf. This would require a small patch to SkPDF. That's much more tightly scoped and thus much less risky. On the flip side, it ties this to tagged PDF generation, might cause problems for bidi text or the like, and would probably result in ActualText being used more than it actually needs to be.

(In reply to James Teh [:Jamie] from comment #7)

Calixte, I'm seeing aria-owns in the math subtree generated by my WIP with something like:
data:text/html,<math><msqrt><mi>x
This breaks Windows screen readers, which can't handle aria-owns in math subtrees as I explained in bug 1998046 comment 9. I thought bug 1998046 was fixed... but was it perhaps only fixed for the AF method? Would fixing this for SE be infeasible?

Can you share the pdf you generated ?

Flags: needinfo?(cdenizet)
Attached file Simple output generated by my WIP. (obsolete) —

Generated from the following markup:
data:text/html,<p>Simple: <math><msqrt><mi>x</mi></msqrt></math></p><p>More complex (exercises ToUnicode fix): <math><mfrac><msqrt><mi>x</mi></msqrt><mrow><mn>2</mn><mi>i</mi></mrow></mfrac></math></p>
Loading this in pdf.js, I see the following raw markup from the MathML subtree for the simple case:
<math><msqrt><mi><span role="none" aria-owns="p2R_mc1"></span></mi></msqrt></math>

Flags: needinfo?(cdenizet)
Depends on: 2055767

We "steal" the text from the text layer only for MathML elements.
Here for <msqrt><mi>x</mi></msqrt> we get:

 > Type: /StructElement
    S: /msqrt
    K:
       > Type: /StructElement
          S: /mi
          K:
             > Type: /StructElement
                S: /NonStruct
                K: 0

So /msqrt and /mi are recognized as MathML elements but /NonStruct is mapped on a span (so not a MathML element) with a role none.
I think, ideally, the /NonStruct element shouldn't exist, I mean we should have S: /mi; K: 0 but if it's too complex I can hack something for pdf.js but unfortunately there's a risk that other viewers will be unhappy with this construction (/NonStruct isn't a part of the MathML ns).

Aha, thank you. No, I'll need to fix this in Gecko, but I don't quite know how yet, since the NonStruct represents the text leaf in our accessibility tree.

Edit: I found a solution. It's in my WIP.

:jamie

do you know if SE is as well supported as AF, or do some readers only support AF?

Until last month Acrobat only supported SE (basically because it didn't understand this at all, so AF got dropped as some unknown key but Structure elements got passed on to AT even if Acrobat didn't know what they were, but the last release added AF support, So currently as far as I know the three pdf readers that understand the mathml tagging at all are Firefox, Acrobat and Foxit and they all support both AF and SE now.

An alternative could be to specify the ActualText in the struct element for a text leaf node

If I understand what you mean there that doesn't work as Actualtext replaces the element. If you have

<msup><mi ActualText=x>x</mi><mn>2</mn></msup>

that will end up being exposed as

<msup>x<mn>2</mn></msup>

which is invalid mathml and generate errors downstream.

You can of course use ActualText on the marked content inside the structure element.

If you want to see how lualatex and Microsoft Word handle tagging radicals, there are some examples at

https://texlive.net/tests/MathML/

(In reply to David Carlisle from comment #13)

An alternative could be to specify the ActualText in the struct element for a text leaf node

If I understand what you mean there that doesn't work as Actualtext replaces the element.

That's not what I meant, but that's because I didn't explain it correctly. :) Gecko would specify ActualText in the struct tree data structure because that's the easiest place we can get at it. However, SkPDF would then expose that as ActualText on the marked content in the content stream.

That said, it's good to know that ActualText on a structure element replaces the element itself, including its semantics. That seems kinda odd to me, but it'd definitely be problematic for this case as you say.

This includes the fix for the missing Unicode characters as well as the removal of the spurious NonStruct elements from the MathML subtree.

Attachment #9610009 - Attachment is obsolete: true

That said, it's good to know that ActualText on a structure element replaces the element itself, including its semantics. That seems kinda odd to me,

I think the model they had in mind was things like text that's technically included as an image to get some funky styling, you can use

/Figure ActualText=hello .. graphic operators for the image

and it gets exposed as pure text hello and not announced as a graphic at all.

In looking at the last generated pdf I noticed several things:

  • for page 1, we have 6 fonts: 4 subsets of CambriaMath and 2 subsets of TimesNewRomanPSMT, it'd be better (in term of size and parsing time) to just have 2 fonts
  • I just checked the ToUnicode of the 1st subset of CambriaMath which is used to render \sqrt{x} and it looks ok
  • in the struct tree, "Simple" (MCID 0) is in > S: /P; K > { S: /NonStruct; K: 0 }. Could we just remove this NonStruct ? it doesn't bring any useful information (in term of semantics), it increases the size of the pdf and the time to parse the struct tree.
Flags: needinfo?(cdenizet)

(In reply to Calixte Denizet (:calixte) from comment #18)

  • for page 1, we have 6 fonts: 4 subsets of CambriaMath and 2 subsets of TimesNewRomanPSMT, it'd be better (in term of size and parsing time) to just have 2 fonts

Way beyond my expertise unfortunately.

  • in the struct tree, "Simple" (MCID 0) is in > S: /P; K > { S: /NonStruct; K: 0 }. Could we just remove this NonStruct ?

This isn't specific to documents containing math. That said, the NonStruct is Gecko's text leaf node. I could remove it as I have done for math, but I worry about cases where a structure element includes multiple text leaves or a mix of text leaves and other structure elements. If you did something like this:

<p>a <b>b</b> c</p>

The bolded "b" would be a separate piece of marked content, but it wouldn't have its own struct element. I'm not sure whether that is ideal or not for readers.

(In reply to Calixte Denizet (:calixte) from comment #18)

  • for page 1, we have 6 fonts: 4 subsets of CambriaMath and 2 subsets of TimesNewRomanPSMT, it'd be better (in term of size and parsing time) to just have 2 fonts

Hmm, I wonder whether my patches have regressed this? Do you get the same result if you just export the same markup to PDF with Nightly?

(In reply to James Teh [:Jamie] from comment #20)

Hmm, I wonder whether my patches have regressed this? Do you get the same result if you just export the same markup to PDF with Nightly?

I tried: data:text/html,<p>Simple: <math><msqrt><mi>x</mi></msqrt></math></p><p>More complex (exercises ToUnicode fix): <math><mfrac><msqrt><mi>x</mi></msqrt><mrow><mn>2</mn><mi>i</mi></mrow></mfrac></math></p> in nightly and I get 1 subset of TimesNewRomanPSMT and 4 of CambriaMath.

This is looking pretty encouraging!

One thing I noticed in your WIP is that nsTextFrame::PropertyProvider::GetToUnicodeText returns the transformed text from a transformed textrun. I think that's probably not desirable; better to return the underlying (untransformed) source text. According to the spec for text-transform,

This property transforms text for styling purposes. It has no effect on the underlying content, and must not affect the content of a plain text copy & paste operation.

Given that it must not affect the content of a plain text copy & paste operation, I think the Unicode text that ends up getting extracted from a PDF should similarly be unaffected by text-transform, and instead reflect the underlying data.

(But it's probably still a good idea to check for a transformed run that is for a masked password, and refuse to expose the original text in this case.)

(In reply to Calixte Denizet (:calixte) from comment #3)

:jfkthame, I tried on Windows and I've the same issue with CambriaMath but only when the glyphs are used as super/sub-scripts.
I don't know how SkPDF builds the ToUnicode map but maybe it's just wrong when the glyphs are from ssty and then it should fallback on the glyph which has induced the substitution.

Yeah, it looks like in SkCairoFTTypeface::getGlyphToUnicodeMap, we basically iterate over the font's cmap and record a mapping back to Unicode for the default glyph that each character maps to. But this won't record any mapping for glyphs that aren't reached directly from the cmap but only through OpenType substitutions. So glyphs that are the result of applying ssty (or other substitution features) will not get a Unicode character mapping, and /ToUnicode won't know about them.

(In reply to Jonathan Kew [:jfkthame] from comment #23)

This also implies that the issue isn't specific to MathML at all, and will affect other content where OpenType substitutions happen. Sure enough, I can reproduce a similar failure with an example like:

data:text/html,<p style="font:25px STIX Two Text">Hello <span style="font-variant:small-caps">World

Saving this to PDF, it looks fine visually, but if I then copy/paste the text from the PDF, I just get "Hello W", because the remaining letters "orld" are substituted with small-caps glyphs, and don't get a /ToUnicode mapping.

Jonathan, any ideas on why we've the same base font subsetted several times (see https://bugzilla.mozilla.org/show_bug.cgi?id=2055405#c21 for an example) ?

Flags: needinfo?(jfkthame)

Not offhand, I'm afraid.... I think that would require digging further into Skia's PDF backend.

(I guess one thing to check would be whether we're using the same Moz2d ScaledFont for all the characters we're outputting, or if the split is happening at a higher level so that Skia would be kinda justified in thinking they should be separate resources.)

Flags: needinfo?(jfkthame)

(In reply to Calixte Denizet (:calixte) from comment #18)

  • in the struct tree, "Simple" (MCID 0) is in > S: /P; K > { S: /NonStruct; K: 0 }. Could we just remove this NonStruct ?

FWIW, I tried this. It kinda works, but it can break if an element contains a mix of text and element children where the logical order of the children differs from the order in which the text is rendered. For example:

data:text/html,<p aria-owns="b c"><span id="b">b</span> <span id="c">c</span> a

The logical order is a b c, but with this change, it will be rendered as b c a. This is because SkPDF has no way of anchoring the text to the correct siblings, so it tries to guess. Most of the time, that guess will probably be correct, but in obscure cases, it won't be. The test case above is very contrived and I doubt it would ever happen in the wild, but there are many ways the rendering order can differ from the logical order and I think we'd run into these eventually.

A better fix would be to tweak SkPDF to accept a placeholder node in the struct tree which can be replaced with the text when it is encountered, rather than the text becoming a child of that node. I suspect that'd be a pretty big lift, though.

I don't think this is a problem for MathML because I don't think MathML elements can accept a mix of text and elements as children.

I don't think this is a problem for MathML because I don't think MathML elements can accept a mix of text and elements as children.

mostly that is true except for token elements (notably mtext).
Basically in mathml-core any element that takes text can take HTML flow content so

<math>
<mfrac>
  <mtext>this <b>bold</b> <a href="https://example.com">link</a></mtext>
  <mtext>over this</mtext>
</mfrac>
</math>

although I'm not sure I really followed all the discussion why NonStruct is being used (or not used)

(In reply to David Carlisle from comment #28)

mostly that is true except for token elements (notably mtext).
Basically in mathml-core any element that takes text can take HTML flow content

Thanks for flagging that. I had thought about it, but haven't addressed it yet. I think I also need to reset the namespace from MathML back to the PDF standard namespace for mtext children. I'm not quite sure how to do that, but I'm guessing explicitly specifying http://iso.org/pdf/ssn will do the trick?

but I'm guessing explicitly specifying http://iso.org/pdf/ssn will do the trick?

Mostly you need the pdf2 namespace http://iso.org/pdf2/ssn

The definition of MathML Structure elements in the PDF 2 spec is (to put it as politely as I can) hopelessly vague.

However for practical reasons it's best to limit what you put in to the pdf to a subset of what is allowed in the html+mathml source.
In particular nvda knows how to read standard pdf structure elements but once it sees /math it hands it all over to MathCat that can't call back
so nested elements only work in as far as MathCat implements them (which it doesn't currently but important ones like a links are on its agenda)
If embedding MathML via AF then it's best to use HTML elements inside mtext (so the attached MathML is normal MathML-Core. But if using MathML SE it's best to use a limited subset of standard structure elements, the best practice guide referenced above says:

Inclusion of PDF structure elements into MathML

MathML (both MathML3 and MathML4) permits the use of
Phrasing
HTML5 tags in any leaf nodes of the MathML structure tree. It is
permitted to use a limited subset of PDF tags as children of structure
elements defined in the MathML namespace.

The table specifies a set of PDF
structure elements permitted as children of the MathML namespace's
mtext structure element in PDF, together with their HTML5
equivalents. The latter are strongly recommended to be used if the
same MathML syntax is expressed via MathML contained in an Associated
File (AF).

Permitted standard PDF structure elements inside MathML

PDF structure element HTML5
Reference,Link a tag
Strong strong tag
Code code tag
Em em tag
Span span tag
Lbl property intent=":equation-label"

As some processors cannot handle PDF structure elements (such as
theLbl structure element) as a child of a MathML structure
element, it is recommended to structure the equation with a label as a
MathMLmtable structure element and represent the label as a MathML
structure elementmtd with Attributeintent equal to
:equation-label.

See Also: → 2056451

Redirect a needinfo that is pending on an inactive user to the triage owner.
:morgan, since the bug has recent activity, could you have a look please?

For more information, please visit BugBot documentation.

Flags: needinfo?(d.p.carlisle) → needinfo?(mreschenberg)

sorry I appear to have switched account and made the last comment as davidc@nag.co.uk rather than d.p.carlisle@gmail.com. so the needinfo flag didn't clear.

Flags: needinfo?(mreschenberg)

(In reply to Jonathan Kew [:jfkthame] from comment #22)

One thing I noticed in your WIP is that nsTextFrame::PropertyProvider::GetToUnicodeText returns the transformed text from a transformed textrun. I think that's probably not desirable; better to return the underlying (untransformed) source text.

It turns out that we were already doing that because we were using mString on the transformed textrun, which is the original text before transformation. But that's also kinda pointless; we may as well just use the same code path (skip chars iterator) regardless of whether the textrun is transformed, except for password masking. I've cleaned this up and am doing a bunch of other cleanup as well.

We need this in order to support tagging MathML in PDF, which requires that we specify the MathML namespace.

Assignee: nobody → jteh
Status: NEW → ASSIGNED

gfx::GlyphBuffer now has optional per-glyph source text and cluster fields, and DrawTarget has a new SupportsGlyphSourceText() method.
Factory::CreateDrawTargetWithSkCanvas now takes a new argument which specifies whether the target supports this.
PrintTargetSkPDF sets this for the PDF page canvas.
When this is set and a caller supplies source text/cluster data, DrawTargetSkia uses SkTextBlobBuilder::allocRunTextPos instead of allocRunPos, falling back to the existing behaviour otherwise.
When splitting glyphs into batches, it avoids splitting glyphs which share a cluster, since each batch only includes the source text for its own glyphs.
Nothing populates the new fields yet, so this is not a behaviour change on its own.

PropertyProvider has a new GetToUnicodeText virtual method, which fetches the source text a range of the textrun was shaped from, as exactly one UTF-16 code unit per textrun character.
The default implementation returns false.
nsTextFrame::PropertyProvider implements it by mapping back to the original DOM text via the skip-chars iterator, even for a transformed textrun, since text-transform must not affect a plain text copy.
The DOM text is never masked, so for a transformed textrun, it explicitly returns false for masked password characters.
Not yet called from anywhere.

gfxTextRun::Draw determines once per run whether to collect source text, based on whether there is a provider and whether DrawTarget::SupportsGlyphSourceText() says the target might be able to use it.
If so, gfxTextRun::DrawGlyphs fetches the source text for the range being drawn from the provider, converts it to UTF-8 along with a mapping from textrun characters to byte offsets, and passes these to gfxFont via TextRunDrawParams, just as it already does for spacing.
The conversion to UTF-8 happens here rather than in the provider because that is what Skia expects (as does cairo's equivalent API), whereas layout deals in UTF-16.
gfxFont::DrawGlyphs maps each glyph to the source text of the character it was drawn for, plus any following characters that have no glyphs of their own.
For example, a ligature glyph maps to the text of all the characters it replaced.
GlyphBufferAzure records a cluster offset for each glyph it buffers and passes the relevant slice of the source text to the DrawTarget whenever it flushes, which can happen part way through a run.
This fixes incorrect /ToUnicode mapping in PDF output for glyphs reached via OpenType substitution, such as MathML sub/superscript styling, when printing directly to SkPDF.

RecordedDrawGlyphs, and the RecordedFillGlyphs/RecordedStrokeGlyphs events derived from it, now take the whole GlyphBuffer rather than just a glyph array and count.
They serialize the buffer's optional source text/cluster fields when present.
Because recordings can come from a less privileged process, the clusters are validated when an event is read, before they can be used to index into the text.
DrawTargetRecording::SupportsGlyphSourceText defers to its recorder, since the target the recording will eventually be replayed against isn't known at record time.
Recorders opt in via DrawEventRecorderPrivate::SetSupportsGlyphSourceText.
This is done for the recorder used for printing in a content process, and for CrossProcessPaint when painting out-of-process iframes for printing, since those are replayed directly into the print target.
Other recording consumers such as remote canvas and WebRender blob images don't opt in, so they don't pay for this.
This is required for the fix in the previous patch to take effect when printing a document in a content process, since content process drawing is recorded and only replayed against the real target (e.g. a tagged SkPDF document) in the parent process.

When PDF outline generation was implemented in bug 1950656, source text wasn't passed to SkPDF.
That meant that SkPDF couldn't accumulate heading names for the outline from the source text.
To work around that, I tweaked Skia slightly to optionally use the alt text as the heading name without exposing the name as the alt attribute in the struct tree.

Now that we pass source text to SkPDF thanks to the earlier patches in this stack, this previous support is redundant and causes duplicate text.
Instead, in most cases, we now allow SkPDF to accumulate heading names from the source text.
Where a heading's name is overridden (e.g. using aria-label), we continue to pass it as alt text.
However, we don't want to expose both alt text and source text, so this tweaks Skia to avoid accumulating heading names from the source text if alt text is explicitly specified for a heading.

Attachment #9649574 - Attachment description: Bug 2055405 part 3: Don't associate nsMathMLChar (square root, etc.) with an accessibility node. r?#layout-reviewers!,#accessibility-platform-reviewers! → Bug 2055405 part 3: Don't associate nsMathMLChar with an accessibility node if the character is implicitly drawn for the element (e.g. square root). r?#layout-reviewers!,#accessibility-platform-reviewers!
Blocks: 2076626
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: