SSL.com: Incorrect Open MPIC Lambda implementation by EJBCA ACME Service
Categories
(CA Program :: CA Certificate Compliance, task)
Tracking
(Not tracked)
People
(Reporter: secauditor, Assigned: secauditor)
Details
(Whiteboard: [ca-compliance] [dv-misissuance])
Preliminary Incident Report
Summary
Incident description: An incorrect Open MPIC Lambda implementation by the EJBCA ACME service allowed DCV to be completed based only on the remote Network Perspectives.
Relevant policies: This is a violation of the following Baseline Requirements:
Section 3.2.2.9 Multi‑Perspective Issuance Corroboration: Multi‐Perspective Issuance Corroboration attempts to corroborate the determinations (i.e., domain validation pass/fail, CAA permission/prohibition) made by the Primary Network Perspective from multiple remote Network Perspectives before Certificate issuance.
Source of incident disclosure: Third Party Reported. Based on a notification from a security reporter received on 2026-04-02 at 01:13 UTC.
SSL.com engineers verified EJBCA behavior as reported by the security reporter. An emergency patch was developed and deployed to remediate the issue, and the vendor has been notified.
The scope of the affected non-expired non-revoked certificates is approximately 1.7 million ACME DV (exact number to be provided in the full incident report). This incident triggered execution of the SSL.com’s Mass Revocation Plan, and we successfully revoked all affected certificates within 24 hours.
SSL.com will post our Full Incident Report on or before 2026-04-17.
Updated•5 months ago
|
Updated•5 months ago
|
Full Incident Report
Summary
-
CA Owner CCADB unique ID: A002038
-
Incident description: An incorrect Open MPIC Lambda implementation by the EJBCA ACME service allowed DCV to be completed based only on the remote Network Perspectives.
-
Timeline summary:
-
Non-compliance start date: 2025-03-13 (date of MPIC enforcement in the SSL.com EJBCA installations)
-
Non-compliance identified date: 2026-04-02
-
Non-compliance end date: 2026-04-02
-
-
Relevant policies: This is a violation of "Section 3.2.2.9 Multi‑Perspective Issuance Corroboration" of the Baseline Requirements: Multi‐Perspective Issuance Corroboration attempts to corroborate the determinations (i.e., domain validation pass/fail, CAA permission/prohibition) made by the Primary Network Perspective from multiple remote Network Perspectives before Certificate issuance.
-
Source of incident disclosure: Third Party Reported. Per investigation triggered by a notification received from a security reporter on 2026-04-02 at 01:13 UTC regarding a possible implementation issue in the EJBCA CA software.
Impact
-
Total number of certificates: 78,443,492, of which 1,770,931 were unexpired, unrevoked certificates
-
Total number of "remaining valid" certificates: Zero (0); all unexpired, unrevoked certificates were revoked within the required 24h timeline.
-
Affected certificate types: This incident affected DV certificates issued through ACME since the enforcement of MPIC in our EJBCA installations.
-
Incident heuristic: The full corpus of unexpired unrevoked affected certificates is disclosed in the Appendix.
-
Was issuance stopped in response to this incident, and why or why not?: No. As described in the incident timeline, the issue was fixed within an hour of confirmation with an emergency patch of the EJBCA software by our CA engineering team.
-
Analysis: Not applicable.
-
Additional considerations: None.
Timeline
-
2025-01-30 Keyfactor releases EJBCA 9.2.0 with MPIC support via OpenMPIC
-
2025-02-12 Internal testing of EJBCA 9.2.0 begins
-
2025-02-24 to 2025-03-13 Release of EJBCA 9.2.1 and 9.2.2. MPIC related fixes based on EJBCA's documentation and internal testing
-
2025-03-13 19:25:49 CA nodes start using MPIC for certificates issued via EJBCA's ACME
-
2025-03-15 Initial MPIC requirements come in effect
-
2026-04-02 01:13:49 A report is received about a potential compliance issue with EJBCA's ACME service
-
2026-04-02 13:44:01 CA team begins investigating the report
-
2026-04-02 16:05:58 Compliance issue is confirmed by CA team. The offending code is identified, and an emergency patch is developed. The issue is tracked by the Compliance Business Unit according to the SSL.com Incident Management Policy. Based on an early evaluation of the affected population, the Mass Revocation Plan is triggered.
-
2026-04-02 16:59:37 All affected ACME systems have been patched
-
2026-04-02 21:15:22 Affected customers are notified
-
2026-04-03 12:42:50 Mass revocation of remaining active certificates initiated
-
2026-04-03 15:57:36 Mass revocation completed and updated OCSP responses & CRLs published
-
2026-04-03 20:08 Preliminary incident report filed to Bugzilla. Root Stores notified.
Related Incidents
| Bug | Date | Description |
|---|---|---|
| 2029643 | 2026-04-06 | Same CA vendor used for implementation of MPIC |
Root Cause Analysis
Contributing Factor #1:
Testing was centered around validating the behavior of the MPIC implementation as outlined in EJBCA's release notes, specifically focusing on ensuring the enforcement of remote perspective validation. This involved integrating EJBCA's MPIC service with our remote perspective infrastructure.
-
Description: An incorrect Open MPIC Lambda implementation by the EJBCA ACME service allowed DCV to be completed based only on the remote Network Perspectives.
-
Timeline:
-
2025-01-30 Keyfactor releases EJBCA 9.2.0 with MPIC support via OpenMPIC
-
2025-02-12 Internal testing of EJBCA 9.2.0 begins
-
2025-02-24 to 2025-03-13 Release of EJBCA 9.2.1 and 9.2.2. MPIC related fixes based on EJBCA's documentation and internal testing
-
2025-03-13 19:25:49 CA nodes start using MPIC for certificates issued via EJBCA's ACME
-
-
Detection: A report was received about a potential compliance issue with EJBCA's ACME service and how it performs its DCV logic when enabling MPIC that was not reported or documented in vendor's release notes.
-
Interaction with other factors: N/A
-
Root Cause Analysis methodology used: 5-Whys
Lessons Learned
-
What went well:
-
The CA Engineering team quickly identified and remediated the issue within EJBCA, leveraging deep knowledge of the underlying thirdparty platform, its architecture, and codebase.
-
Rapid response based on resolute decision-making and close coordination between teams in the execution of a well-documented Mass Revocation Plan ensuring timely remediation.
-
-
What didn't go well:
- Our test suites did not include an undocumented use case by EJBCA. We were not aware of a testing coverage gap based on information provided by the third-party vendor.
-
Where we got lucky:
- The vast majority of affected certificates had a validity period of 90 days or less, which resulted in a small percentage of the active affected certificates at time of discovery.
-
Additional: None
Action Items
| Action Item | Kind | Corresponding Root Cause(s) | Evaluation Criteria | Due Date | Status |
|---|---|---|---|---|---|
| 1. Deploy a fix to the EJBCA ACME service | Prevent | Root Cause # 1 | Compliance testing to confirm DCV by the Primary Perspective | 2026-04-02 | Completed |
| 2. Update testing procedures to include explicit positive and negative test cases that distinguish between Primary Perspective and Remote Perspective failure scenarios and incorporate DCV Inspector validation. | Prevent | Root Cause # 1 | Test coverage confirmed through internal change management procedures. | 2026-04-30 | Open |
Appendix
Attached
Due to the large size of the Appendix, it can be found here: https://legal.ssl.com/documents/appendixsha256.txt
Comment 3•5 months ago
|
||
Thank you for the detailed incident report.
To ensure that our understanding of the MPIC enforcement model aligns correctly with the Baseline Requirements, I would like to ask for a technical clarification regarding the role of the Primary Network Perspective.
BR Section 3.2.2.9 defines Multi-Perspective Issuance Corroboration as a mechanism to corroborate DCV and CAA determinations made by the Primary Network Perspective.
In this incident, was the Primary Network Perspective performing DCV and CAA validation and making the authoritative pass/fail determinations as intended, with issuance incorrectly proceeding based solely on corroboration results from remote Network Perspectives due to an enforcement or gating issue?
Or, alternatively, was DCV effectively completed without the Primary Network Perspective’s validation logic being executed, with remote Network Perspectives independently performing the validation?
Clarifying where the DCV decision point resided in the implementation would be very helpful for broader industry understanding of correct MPIC architectures.
One small clarification regarding the Action Items table:
I noticed that Action Item #2 is marked as "Open".
Per the CCADB Incident Reporting Guidelines, Action Item status values are typically expressed as one of "Ongoing", "Complete", "Delayed", or "Canceled".
Would "Ongoing" be the intended status here?
Comment 4•5 months ago
|
||
Hi,
In light of Self-audits required by the TLS Baseline requirements, did you consider that in your root cause analysis?
Given the number of certificates affected by this incident, were any of those included in the sample set of your 3% self-audits, so that the issue could have been detected earlier than it eventually was discovered?
Wouldn't it be one of the things in self-audit to review that MPIC is performed in accordance with the requirements?
Thank you to SSL.com for filing this incident report. That said, the report reads more as a status update than a root cause analysis — it is notably light on self-reflection and contributing factors. The community would benefit from SSL.com examining not just what went wrong technically, but what organizational, process, and oversight failures allowed this condition to exist and persist undetected. The questions below are offered in that spirit.
1. Issuance was not halted upon receiving the report
The report states that a security reporter notified SSL.com of a potential validation failure on 2026-04-02 at 01:13 UTC. Issuance does not appear to have been halted at that point.
- At 01:13 UTC, what was SSL.com's understanding of the severity of the reported issue, and on what basis was a decision made — explicit or implicit — to allow continued issuance?
- If the decision was that issuance could safely continue while a potential DCV compliance failure was being assessed, who made that decision, under what authority, and what criteria governed it?
- If no such decision was made and issuance simply continued by default, does SSL.com consider that an acceptable posture for a CA operating under root program requirements? Does this reflect documented policy, or the absence of one?
- How many certificates were issued in the window between receipt of the report and the point at which issuance was halted or remediated? Please provide exact counts by issuance path.
- Does the team responsible for triaging problem reports have the authority and operational capability to immediately suspend a specific issuance path — at the level of a validation method, request protocol, or third-party component — without escalation or a full system shutdown, and was that capability available and considered during this incident?
Allowing issuance to continue while a potential misissuance condition is unresolved is itself a concern independent of the underlying defect. It would help the community understand what SSL.com's incident response policy actually requires in this scenario, and whether it was followed.
2. Twelve-hour delay before investigation began, and revocation clock interpretation
Timeline as reported:
- 2026-04-02 01:13 UTC — problem report received
- 2026-04-02 13:44 UTC — investigation started
- 2026-04-02 16:05 UTC — compliance issue confirmed by CA team
- 2026-04-03 12:42 UTC — mass revocation initiated
- 2026-04-03 15:57 UTC — mass revocation completed
This is an approximately 12.5-hour gap between report receipt and investigation start, and approximately 35 hours 29 minutes between report receipt and revocation being initiated.
On the investigation delay:
- Was this report received through a monitored channel at 01:13 UTC? If so, what caused the delay in triage until 13:44 UTC?
- If the report arrived outside of staffed hours, does SSL.com's 24×7 commitment extend to incident triage and initial investigation, or only to monitoring? What does "24×7" mean operationally for SSL.com's incident response program?
- Was the delay caused by a staffing gap, a process failure, a prioritization decision, or something else?
While the BR 4.9.5 investigation trigger allows up to 24 hours, Mozilla's own incident response guidance is explicit that in misissuance cases a CA should almost always immediately cease issuance from the affected part of its PKI. This was not a marginal or ambiguous compliance question — it was a reported DCV control failure. A 12.5-hour delay before investigation even began warrants a specific explanation of why that is consistent with treating this incident with the urgency it warranted.
On the revocation clock:
BR 4.9.5 is explicit: "the period from receipt of the Certificate Problem Report or revocation-related notice to published revocation MUST NOT exceed the time frame set forth in Section 4.9.1.1." The clock runs from report receipt — not from internal confirmation of the issue.
This interpretation is not novel. In a prior community discussion involving Entrust (Bug 1520876), it was directly established that "the deadline will be based on the time of notification and not the time the investigation is complete."
From report receipt at 01:13 UTC on 2026-04-02 to revocation being initiated at 12:42 UTC on 2026-04-03 is approximately 35 hours 29 minutes — well outside the 24-hour requirement under this interpretation.
- On what basis does SSL.com interpret the 24-hour revocation clock as running from internal confirmation at 16:05 UTC rather than from report receipt at 01:13 UTC?
- If SSL.com's position is that the clock starts at confirmation, how does that square with the plain language of BR 4.9.5 and the community precedent established in the Entrust discussion?
- Does SSL.com consider its revocation timeline to be compliant with BR 4.9.1.1? If so, please provide the specific basis for that interpretation.
3. Emergency patch developed and deployed to production in under one hour
The report states that after confirming the issue, an emergency patch was developed and deployed to production in approximately one hour.
- Was the fix validated in a staging or pre-production environment prior to production deployment? If not, why not, and does SSL.com have a standing exception process for emergency changes that bypasses staging validation?
- What specific testing was performed in that window? Were there negative tests and boundary conditions — specifically, was the fix verified to correctly enforce primary-perspective corroboration, not merely to reject the previously defective path?
- Was a formal code review conducted prior to deployment? If so, by how many reviewers, and what was the scope relative to the change?
- How did SSL.com weigh the risk of continued misissuance against the risk of introducing new defects through an accelerated deployment, and was that analysis documented?
- Has SSL.com since performed a post-deployment review to confirm the fix behaved correctly in production and did not introduce regressions elsewhere in the issuance pipeline?
One hour from confirmation to production — for a change to certificate validation logic — warrants a detailed account of what quality controls were actually applied.
4. Attribution to a third-party component
SSL.com attributes this defect to an incorrect implementation within EJBCA's ACME service. The Baseline Requirements are clear that a CA is fully responsible for the compliance of its issuance operations regardless of whether components are developed in-house or by vendors.
- What acceptance testing did SSL.com perform on EJBCA's MPIC integration before placing it into production? Was MPIC behavior — including primary-perspective corroboration — part of the acceptance test criteria?
- If EJBCA's MPIC implementation was not independently tested by SSL.com prior to production deployment, what was SSL.com relying on to ensure MPIC compliance other than the correctness of the vendor implementation?
- Given MPIC's significance as a required DV issuance control, was EJBCA's implementation subject to any independent security or compliance review by SSL.com — separate from EJBCA's own testing — before it was used in production issuance?
- Does SSL.com maintain any independent verification controls — downstream of the issuance pipeline — that would detect a case where a required control was not correctly enforced? If so, why did those controls not detect this condition? If not, what compensating mechanism existed?
- Approximately 1.7 million certificates were issued under this defective configuration. Over what time period did this occur, and were there no internal signals — error rates, audit log anomalies, compliance checks — that indicated something was wrong during that period?
The concern is not simply about a vendor bug. It is about whether SSL.com had any independent mechanism to verify that its compliance obligations were being met in practice — and if not, why not.
5. Scope and completeness of the post-incident review
The preliminary report focuses on EJBCA's ACME service, but the community has no visibility into the breadth of SSL.com's issuance infrastructure.
- Has SSL.com audited all other issuance paths to confirm this defect was isolated to this specific component? What was the scope and methodology of that review?
- If that review has not yet been completed, when will it be, and will the full incident report include its findings?
- Were any other integration points found to have inconsistent behavior during the post-incident review?
- This incident was discovered by an external security reporter, not by SSL.com's own monitoring or compliance controls. What does SSL.com believe this indicates about the effectiveness of its existing compliance verification program?
6. Mass revocation execution, ARI support, and preparedness
The questions below are not about whether revocation was appropriate — it was — but about execution quality, subscriber impact, and what SSL.com has learned to improve future preparedness.
ARI and automated renewal:
- Did SSL.com's ACME service support ACME Renewal Information (ARI) at the time of this incident?
- If so, what percentage of affected ACME clients triggered automated reissuance based on the ARI signal prior to revocation being enforced?
- If ARI was not supported, does SSL.com consider that a gap given the scale at which it operates, and is ARI support now planned or in progress?
Subscriber impact:
- Of the approximately 1.7 million revoked certificates, what percentage had already been successfully replaced before revocation was initiated?
- Does SSL.com have visibility into how many subscribers experienced unexpected service disruption as a result of revocation, versus those who had already renewed?
- What communication channels and tooling did SSL.com use to notify affected subscribers, and how quickly were those notifications issued after the decision to revoke was made?
Preparedness and lessons learned:
- Did SSL.com have a tested mass revocation plan in place prior to this incident, and how did actual execution compare to that plan?
- What gaps or unexpected challenges were identified during the execution of this revocation event?
- What specific improvements to SSL.com's mass revocation preparedness, subscriber communication, and ARI support has this incident prompted?
A mass revocation of 1.7 million certificates is an operationally significant event. The community would benefit from understanding not just that SSL.com met the 24-hour requirement, but what this event revealed about their preparedness and what concrete improvements are being made as a result.
(In reply to SECOM Trust Systems - ONO Fumiaki from comment #3)
Thank you for the detailed incident report.
To ensure that our understanding of the MPIC enforcement model aligns correctly with the Baseline Requirements, I would like to ask for a technical clarification regarding the role of the Primary Network Perspective.
BR Section 3.2.2.9 defines Multi-Perspective Issuance Corroboration as a mechanism to corroborate DCV and CAA determinations made by the Primary Network Perspective.
In this incident, was the Primary Network Perspective performing DCV and CAA validation and making the authoritative pass/fail determinations as intended, with issuance incorrectly proceeding based solely on corroboration results from remote Network Perspectives due to an enforcement or gating issue?
Or, alternatively, was DCV effectively completed without the Primary Network Perspective’s validation logic being executed, with remote Network Perspectives independently performing the validation?
It would be the later, in which DCV effectively completed without the Primary Network Perspective’s validation logic being executed, with remote Network Perspectives independently performing the validation (2 remote perspectives from the beginning of MPIC use for EJBCA ACME on March 13, 2025 and 3 remote perspectives starting on March 12, 2026).
Clarifying where the DCV decision point resided in the implementation would be very helpful for broader industry understanding of correct MPIC architectures.
One small clarification regarding the Action Items table:
I noticed that Action Item #2 is marked as "Open".
Per the CCADB Incident Reporting Guidelines, Action Item status values are typically expressed as one of "Ongoing", "Complete", "Delayed", or "Canceled".Would "Ongoing" be the intended status here?
Yes, thank you for pointing this out. The intended status was “ongoing” for Action Item #2.
(In reply to Antti Backman from comment #4)
Hi,
In light of Self-audits required by the TLS Baseline requirements, did you consider that in your root cause analysis?
Given the number of certificates affected by this incident, were any of those included in the sample set of your 3% self-audits, so that the issue could have been detected earlier than it eventually was discovered?
Wouldn't it be one of the things in self-audit to review that MPIC is performed in accordance with the requirements?
The failure of detective controls, and more specifically of the 3% self-audits, is an issue that has been discussed internally as part of the Root Cause Analysis. Our investigation showed that the non-detection of the issue stems from the lack of verbosity in EJBCA’s ACME DCV records. Our self-audit tooling relies on an EJBCA pass/fail record, reflecting the final status of the ACME challenge. We intend to collaborate with our vendor to incorporate database support for storing the underlying DCV/MPIC challenge data within EJBCA, thereby enhancing future self-audit testing capabilities. Given the necessity of involving our third-party CA vendor, we plan to include this commitment in our closing report.
(In reply to trusten from comment #5)
Hello, trusten. Thanks for your detailed questions. We’re currently working on them and should have a response for you by next week.
Comment 9•4 months ago
|
||
(In reply to SSL.com from comment #6)
Thank you for your response and clarification.
| Assignee | ||
Comment 10•4 months ago
|
||
(In reply to trusten from comment #5)
Thank you to SSL.com for filing this incident report. That said, the report reads more as a status update than a root cause analysis — it is notably light on self-reflection and contributing factors. The community would benefit from SSL.com examining not just what went wrong technically, but what organizational, process, and oversight failures allowed this condition to exist and persist undetected. The questions below are offered in that spirit.
1. Issuance was not halted upon receiving the report
The report states that a security reporter notified SSL.com of a potential validation failure on 2026-04-02 at 01:13 UTC. Issuance does not appear to have been halted at that point.
- At 01:13 UTC, what was SSL.com's understanding of the severity of the reported issue, and on what basis was a decision made — explicit or implicit — to allow continued issuance?
- If the decision was that issuance could safely continue while a potential DCV compliance failure was being assessed, who made that decision, under what authority, and what criteria governed it?
- If no such decision was made and issuance simply continued by default, does SSL.com consider that an acceptable posture for a CA operating under root program requirements? Does this reflect documented policy, or the absence of one?
- How many certificates were issued in the window between receipt of the report and the point at which issuance was halted or remediated? Please provide exact counts by issuance path.
- Does the team responsible for triaging problem reports have the authority and operational capability to immediately suspend a specific issuance path — at the level of a validation method, request protocol, or third-party component — without escalation or a full system shutdown, and was that capability available and considered during this incident?
Allowing issuance to continue while a potential misissuance condition is unresolved is itself a concern independent of the underlying defect. It would help the community understand what SSL.com's incident response policy actually requires in this scenario, and whether it was followed.
We confirm that the SSL.com Incident Management Policy (IMP) was followed throughout the incident.
In particular, with regards to the issue at hand:
A report was received at 01:13 UTC about a potential compliance issue with EJBCA’s ACME service. To clarify, the report was not a Certificate Problem Report or a notification of an observed SSL.com violation. It was a report by a security researcher regarding a potential bug in EJBCA’s ACME implementation that might affect CAs that rely on it.
The report triggered an investigation by the SSL.com engineering team to review the code, test and confirm the bug, and its implications to the SSL.com systems. An emergency patch was created as part of this code review and testing process. This was feasible because of the very nature of the issue: a plain if-then-else application logic mistake which did not allow execution of the standard DCV code block when EJBCA’s ACME MPIC feature is enabled.
The Compliance Business Unit (CBU) received an internal notification informing them about the verification of the issue and the availability of an emergency fix. All affected servers were patched shortly after the approval of the emergency fix as reported in the timeline.
For the record, the volume of TLS certificates issued through ACME from the reporting of EJBCA’s potential problematic behavior to the remediation of the issue in all production services is 48,749.
Process wise:
The incident management process is governed by the IMP which stipulates suspected compliance violations (Security Events) must be investigated, and a determination of the violation against the specific requirement must be made to escalate to a confirmed violation (Security Incident) and act accordingly.
Triaging, and the entire Incident Management process, is controlled by the CBU per mandate by the IMP and the Policy Management Authority (PMA). The CBU has the authority to request immediate suspension of an issuance path (e.g. ACME), subject to approval by the leadership (Executive Leadership Council). Execution of the final order requires the involvement of the relevant operational teams, i.e. there is no Emergency Stop Switch.
At this point, it is worth noting that upon verification of the violation, in case the issue needs to be contained (e.g., to avoid further mis-issuances), suspension of the issuance path is required by the IMP. An exception may be granted such as in cases like this incident, that the remediation of the issue was immediately available.
2. Twelve-hour delay before investigation began, and revocation clock interpretation
Timeline as reported:
- 2026-04-02 01:13 UTC — problem report received
- 2026-04-02 13:44 UTC — investigation started
- 2026-04-02 16:05 UTC — compliance issue confirmed by CA team
- 2026-04-03 12:42 UTC — mass revocation initiated
- 2026-04-03 15:57 UTC — mass revocation completed
This is an approximately 12.5-hour gap between report receipt and investigation start, and approximately 35 hours 29 minutes between report receipt and revocation being initiated.
On the investigation delay:
- Was this report received through a monitored channel at 01:13 UTC? If so, what caused the delay in triage until 13:44 UTC?
- If the report arrived outside of staffed hours, does SSL.com's 24×7 commitment extend to incident triage and initial investigation, or only to monitoring? What does "24×7" mean operationally for SSL.com's incident response program?
- Was the delay caused by a staffing gap, a process failure, a prioritization decision, or something else?
While the BR 4.9.5 investigation trigger allows up to 24 hours, Mozilla's own incident response guidance is explicit that in misissuance cases a CA should almost always immediately cease issuance from the affected part of its PKI. This was not a marginal or ambiguous compliance question — it was a reported DCV control failure. A 12.5-hour delay before investigation even began warrants a specific explanation of why that is consistent with treating this incident with the urgency it warranted.
On the revocation clock:
BR 4.9.5 is explicit: "the period from receipt of the Certificate Problem Report or revocation-related notice to published revocation MUST NOT exceed the time frame set forth in Section 4.9.1.1." The clock runs from report receipt — not from internal confirmation of the issue.
This interpretation is not novel. In a prior community discussion involving Entrust (Bug 1520876), it was directly established that "the deadline will be based on the time of notification and not the time the investigation is complete."
From report receipt at 01:13 UTC on 2026-04-02 to revocation being initiated at 12:42 UTC on 2026-04-03 is approximately 35 hours 29 minutes — well outside the 24-hour requirement under this interpretation.
- On what basis does SSL.com interpret the 24-hour revocation clock as running from internal confirmation at 16:05 UTC rather than from report receipt at 01:13 UTC?
- If SSL.com's position is that the clock starts at confirmation, how does that square with the plain language of BR 4.9.5 and the community precedent established in the Entrust discussion?
- Does SSL.com consider its revocation timeline to be compliant with BR 4.9.1.1? If so, please provide the specific basis for that interpretation.
As mentioned in question #1, no Certificate Problem Report was submitted. Instead, an email was sent at 2026-04-02 01:13 UTC to an alternate SSL.com address by a security researcher regarding a potential bug in EJBCA’s ACME implementation. The email did not allege, or provide evidence of an SSL.com compliance violation, and it was subsequently routed to our engineering team for investigation. Through this investigation, at 2026-04-02 16:05 UTC, the engineering team confirmed the bug affected our ACME implementation. Upon verification of the violation, the 24-hour revocation clock started in compliance with BR 4.9.1.1.
3. Emergency patch developed and deployed to production in under one hour
The report states that after confirming the issue, an emergency patch was developed and deployed to production in approximately one hour.
- Was the fix validated in a staging or pre-production environment prior to production deployment? If not, why not, and does SSL.com have a standing exception process for emergency changes that bypasses staging validation?
- What specific testing was performed in that window? Were there negative tests and boundary conditions — specifically, was the fix verified to correctly enforce primary-perspective corroboration, not merely to reject the previously defective path?
- Was a formal code review conducted prior to deployment? If so, by how many reviewers, and what was the scope relative to the change?
- How did SSL.com weigh the risk of continued misissuance against the risk of introducing new defects through an accelerated deployment, and was that analysis documented?
- Has SSL.com since performed a post-deployment review to confirm the fix behaved correctly in production and did not introduce regressions elsewhere in the issuance pipeline?
One hour from confirmation to production — for a change to certificate validation logic — warrants a detailed account of what quality controls were actually applied.
Yes, SSL.com has a standing emergency change process that was invoked here that allows for the bypass of the normal change process that would’ve included the staging environment. The fix was validated in the dev environment, and code review was performed by a separate engineer. The emergency change process was used here given the simplicity of the nature of the fix as well as the urgency to fix the issue.
The tests in the dev environment included verifying that there were validation attempts from all perspectives (primary and secondary), that failure of either one leads to the challenge not being satisfied and that both being satisfied leads to the challenge being satisfied.
The risk with this particular fix was determined to be low due to the simplicity of the fix once it was determined what the issue was. This analysis was made by subject matter experts that were handling the incident. Additionally, given the straightforward nature of the fix, the post-deployment review was conducted immediately to determine that it was performing as intended and as it did in the dev environment testing, however that was not formally documented at the time it occurred.
4. Attribution to a third-party component
SSL.com attributes this defect to an incorrect implementation within EJBCA's ACME service. The Baseline Requirements are clear that a CA is fully responsible for the compliance of its issuance operations regardless of whether components are developed in-house or by vendors.
- What acceptance testing did SSL.com perform on EJBCA's MPIC integration before placing it into production? Was MPIC behavior — including primary-perspective corroboration — part of the acceptance test criteria?
- If EJBCA's MPIC implementation was not independently tested by SSL.com prior to production deployment, what was SSL.com relying on to ensure MPIC compliance other than the correctness of the vendor implementation?
- Given MPIC's significance as a required DV issuance control, was EJBCA's implementation subject to any independent security or compliance review by SSL.com — separate from EJBCA's own testing — before it was used in production issuance?
- Does SSL.com maintain any independent verification controls — downstream of the issuance pipeline — that would detect a case where a required control was not correctly enforced? If so, why did those controls not detect this condition? If not, what compensating mechanism existed?
- Approximately 1.7 million certificates were issued under this defective configuration. Over what time period did this occur, and were there no internal signals — error rates, audit log anomalies, compliance checks — that indicated something was wrong during that period?
The concern is not simply about a vendor bug. It is about whether SSL.com had any independent mechanism to verify that its compliance obligations were being met in practice — and if not, why not.
Acceptance criteria for any change made to SSL.com's PKI system is determined and controlled through our PKI Change Management Policy. Based on our understanding of EJBCA's ACME MPIC feature, as documented, we determined our implementation was compliant and met our criteria for production use by Trusted Role personnel within SSL.com.
SSL.com did have tests in place to check MPIC behaviors (passing correct parameters to OpenMPIC) and also testing of OpenMPIC to operate as expected when provided the appropriate parameters including failed challenge behavior. We did not expand our test coverage as part of EJBCA's update since we had no indication additional testing was required as a result of their change. Our understanding of EJBCA's ACME MPIC behavior was based on their then-current documentation and our observation of successful testing prior to the update. We have expanded our testing regimen when promoting EJBCA code to production and are including specific verifications related to this incident’s use case as per our Action Item #2.
5. Scope and completeness of the post-incident review
The preliminary report focuses on EJBCA's ACME service, but the community has no visibility into the breadth of SSL.com's issuance infrastructure.
- Has SSL.com audited all other issuance paths to confirm this defect was isolated to this specific component? What was the scope and methodology of that review?
- If that review has not yet been completed, when will it be, and will the full incident report include its findings?
- Were any other integration points found to have inconsistent behavior during the post-incident review?
- This incident was discovered by an external security reporter, not by SSL.com's own monitoring or compliance controls. What does SSL.com believe this indicates about the effectiveness of its existing compliance verification program?
Our investigation included all issuance paths and nodes that perform MPIC as part of the domain validation. The only paths found to be affected were the ones that relied on the EJBCA ACME service, i.e. a third-party system.
The issue did not affect any of the in-house systems that utilize the same underlying MPIC infrastructure. We were able to confirm through testing and code review that the MPIC business logic implemented in our in-house systems was correct and according to the standards.
Regarding the failure of self-audits as detective control, please refer to Comment #7.
6. Mass revocation execution, ARI support, and preparedness
The questions below are not about whether revocation was appropriate — it was — but about execution quality, subscriber impact, and what SSL.com has learned to improve future preparedness.
ARI and automated renewal:
- Did SSL.com's ACME service support ACME Renewal Information (ARI) at the time of this incident?
No
- If so, what percentage of affected ACME clients triggered automated reissuance based on the ARI signal prior to revocation being enforced?
N/A as ARI was not available
- If ARI was not supported, does SSL.com consider that a gap given the scale at which it operates, and is ARI support now planned or in progress?
We support ARI now. Although this capability was part of an approved deployment plan, it had not been operational at the time the incident occurred.
Subscriber impact:
- Of the approximately 1.7 million revoked certificates, what percentage had already been successfully replaced before revocation was initiated?
11% had been replaced before revocation was initiated.
- Does SSL.com have visibility into how many subscribers experienced unexpected service disruption as a result of revocation, versus those who had already renewed?
No disruptions were reported to us by this incident.
- What communication channels and tooling did SSL.com use to notify affected subscribers, and how quickly were those notifications issued after the decision to revoke was made?
Subscriber notifications were distributed via email through SSL.com’s CRM notification system.
Preparedness and lessons learned:
- Did SSL.com have a tested mass revocation plan in place prior to this incident, and how did actual execution compare to that plan?
A documented Mass Revocation Plan existed at the time of the incident, and our first test of the plan was scheduled for May.
- What gaps or unexpected challenges were identified during the execution of this revocation event?
During the distribution of bulk email notifications, a limited number of unexpected communication issues were encountered (please see the next answer, which further explains challenges that were faced) and promptly resolved within minutes of the initial send. Aside from the issues noted before, no other gaps or unforeseen challenges were observed during the execution of the revocation.
- What specific improvements to SSL.com's mass revocation preparedness, subscriber communication, and ARI support has this incident prompted?
During the bulk email notification effort, SSL.com identified limitations in the CRM-based email tooling for emergency communications, including message framing and segmentation that were oriented toward routine subscriber messaging. Delivery was also impacted by pre-existing suppression list data (e.g., master suppression, unsubscribe, and bounce lists), which reduced the reachable population and required additional operational handling. Following this incident, SSL.com reviewed communication workflows and updated internal notification procedures for future emergency events.
A mass revocation of 1.7 million certificates is an operationally significant event. The community would benefit from understanding not just that SSL.com met the 24-hour requirement, but what this event revealed about their preparedness and what concrete improvements are being made as a result.
| Assignee | ||
Comment 11•4 months ago
|
||
This is an update to report our progress with remediation actions.
The following action item has been completed:
| Action Item | Kind | Corresponding Root Cause(s) | Evaluation Criteria | Due Date | Status |
|---|---|---|---|---|---|
| Update testing procedures to include explicit positive and negative test cases that distinguish between Primary Perspective and Remote Perspective failure scenarios and incorporate DCV Inspector validation. | Prevent | Root Cause # 1 | Test coverage confirmed through internal change management procedures. | 2026-04-30 | Completed |
SSL.com continues to monitor this bug, and requests for our next update to be set for 2026-05-07.
Updated•4 months ago
|
| Assignee | ||
Comment 12•4 months ago
|
||
SSL.com continues to monitor this bug for any additional questions or responses. We kindly request our next update to be set for 2026-05-14, as we start to prepare our closing summary.
Comment 13•4 months ago
|
||
(In reply to SSL.com from comment #10)
(In reply to trusten from comment #5)
Thank you for the detailed responses. Several points remain unclear and require clarification.
1. Issuance containment
SSL.com has confirmed there is no immediate “Emergency Stop” capability and that suspension requires escalation and coordination. In this incident, issuance continued and 48,749 certificates were issued after report receipt.
Mozilla guidance states that in cases of potential misissuance, a CA should almost always immediately cease issuance from the affected part of its PKI.
- Can SSL.com confirm it will follow this expectation going forward, including before the issue is fully confirmed?
- What changes will be implemented to ensure SSL.com’s processes and systems are capable of meeting this expectation in practice?
2. Problem reporting and intake
This issue was reported via an alternate channel rather than through a formal problem reporting path.
BR 4.9.5 requires CAs to maintain a continuous 24×7 ability to accept and respond to Certificate Problem Reports and revocation-related notices.
- Can SSL.com confirm it will take the necessary measures to ensure security and compliance issues are consistently reported through appropriate and clearly identifiable channels?
- What changes will be made to ensure all incoming reports, regardless of channel, are identified, escalated, and triaged with appropriate urgency?
3. Mass revocation readiness and testing
SSL.com has confirmed that its Mass Revocation Plan had not been tested prior to this incident, despite Mozilla Root Store Policy Section 6.1.3 requiring not only maintaining but also testing such plans.
- What was the rationale for not performing any form of testing, simulation, or validation of the plan in the period following the requirement becoming effective?
- Can SSL.com confirm that mass revocation plan testing is now implemented and will be conducted in a manner that demonstrates operational feasibility, rather than deferring testing within a nominal annual cycle?
Updated•4 months ago
|
| Assignee | ||
Comment 14•4 months ago
|
||
(In reply to trusten from comment #13)
(In reply to SSL.com from comment #10)
(In reply to trusten from comment #5)
Thank you for the detailed responses. Several points remain unclear and require clarification.
1. Issuance containment
SSL.com has confirmed there is no immediate “Emergency Stop” capability and that suspension requires escalation and coordination. In this incident, issuance continued and 48,749 certificates were issued after report receipt.
Mozilla guidance states that in cases of potential misissuance, a CA should almost always immediately cease issuance from the affected part of its PKI.
- Can SSL.com confirm it will follow this expectation going forward, including before the issue is fully confirmed?
- What changes will be implemented to ensure SSL.com’s processes and systems are capable of meeting this expectation in practice?
SSL.com fully adheres to its CP/CPS, the applicable CA/B Forum requirements and the Root Store Policies, including the processing of Certificate Problem Reports and Incident Management. SSL.com processed a report of a potential bug from outside our CPR mechanism without undue delay, tested and completed the initial triaging within the same day. Once the violation was confirmed, SSL.com escalated the issue to an incident and fully complied with the strict revocation and incident reporting timelines. As previously stated, in this instance, an exception was granted to the normal suspension of the issuance path in our IMP given the simplicity of the fix and how quickly it would be able to be in place.
The expectation does not include mandating suspension of certificate issuance based on any suspicion before confirmation as you suggest. This would create an attack surface on CAs where anyone could send an email claiming a software bug that results in shutting down of issuance before confirmation of the issue or its impact.
2. Problem reporting and intake
This issue was reported via an alternate channel rather than through a formal problem reporting path.
BR 4.9.5 requires CAs to maintain a continuous 24×7 ability to accept and respond to Certificate Problem Reports and revocation-related notices.
- Can SSL.com confirm it will take the necessary measures to ensure security and compliance issues are consistently reported through appropriate and clearly identifiable channels?
- What changes will be made to ensure all incoming reports, regardless of channel, are identified, escalated, and triaged with appropriate urgency?
Yes, your prompt is correct. For accuracy, BR 4.9.3 (not 4.9.5) does mandate that Certificate Authorities (CAs) maintain a “continuous 24x7 ability to accept and respond to revocation requests and Certificate Problem Reports.” As a publicly trusted CA, we fully comply with our obligation to keep our CP/CPS and CCADB profiles updated with our current Certificate Problem Reporting (CPR) procedures and links and maintain a capacity to be able to respond accordingly.
To the best of our knowledge, there is no expectation on a CA to control how external parties submit other types of reports to us.
Additionally, as a reminder, in response to your previous questions (comment 10, question 1, and 2), “No Certificate Problem Report was submitted,” this was an email about a potential software bug.
SSL.com continues to encourage reporters who submit reports through alternative channels, if the report falls under the guidelines of a CPR, to utilize our clearly outlined CPR processes, which are available in our policy documents, and in our profile on CCADB.
3. Mass revocation readiness and testing
SSL.com has confirmed that its Mass Revocation Plan had not been tested prior to this incident, despite Mozilla Root Store Policy Section 6.1.3 requiring not only maintaining but also testing such plans.
- What was the rationale for not performing any form of testing, simulation, or validation of the plan in the period following the requirement becoming effective?
- Can SSL.com confirm that mass revocation plan testing is now implemented and will be conducted in a manner that demonstrates operational feasibility, rather than deferring testing within a nominal annual cycle?
As a reminder, while having a formalized Mass Revocation Plan is a new requirement, it is not a new requirement for a CA to be able to revoke any number of certificates within the required timelines.
The adoption of these requirements in our CP/CPS v1.27 in November 2025, set out a 12-month leeway to execute the testing of the plan. Considering SSL.com has already had the ability to meet the revocation expectation prior to the Mass Revocation Plan requirements, and to align with our audit cycle, an earlier execution of the plan was not determined to be necessary.
SSL.com confirms that our mass revocation plan testing is fully compliant with the BRs and the MRSP.
Updated•4 months ago
|
| Assignee | ||
Comment 15•3 months ago
|
||
Report Closure Summary
-
Incident description: An incorrect Open MPIC Lambda implementation by the EJBCA ACME service allowed DCV to be completed based only on the remote Network Perspectives.
-
Incident Root Cause(s): Testing was centered around validating the behavior of the MPIC implementation as outlined in EJBCA’s release notes, specifically focusing on ensuring the enforcement of remote perspective validation. This involved integrating EJBCA’s MPIC service with our remote perspective infrastructure.
-
Remediation description: Deploy a fix to the EJBCA ACME service and update testing procedures to include explicit positive and negative test cases that distinguish between Primary Perspective and Remote Perspective failure scenarios and incorporate DCV Inspector validation.
-
Commitment summary: SSL.com has implemented an internal process to update our testing procedures. These procedures now include explicit positive and negative test cases that differentiate between Primary Perspective and Remote Perspective failure scenarios. Additionally, DCV Inspector validation has been incorporated into the testing process. SSL.com values the feedback and questions received on this bug report and is committed to maintaining compliance with CA/B Forum requirements, root program expectations, and industry best practices.
All Action Items disclosed in this report have been completed as described, and we request its closure.
Comment 16•3 months ago
|
||
This is a final call for comments or questions on this Incident Report.
Otherwise, it will be closed on approximately 2026-05-28.
Updated•3 months ago
|
Description
•