Closed Bug 1677239 Opened 5 years ago Closed 5 years ago

IdenTrust: Service Degradation

Categories

(CA Program :: CA Certificate Compliance, task)

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: roots, Assigned: roots)

Details

(Whiteboard: [ca-compliance] [ocsp-failure])

User Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.193 Safari/537.36

Steps to reproduce:

On November 6, 2020 we were made aware of failed responses to resolve a DNS lookup. This could have caused degradation to some IdenTrust services.
We expect to provide a complete Incident Report no later than November 20, 2020

Assignee: bwilson → roots
Status: UNCONFIRMED → ASSIGNED
Type: enhancement → task
Ever confirmed: true
Whiteboard: [ca-compliance]
  1. How your CA first became aware of the problem (e.g. via a problem report submitted to your Problem Reporting Mechanism, a discussion in mozilla.dev.security.policy, a Bugzilla bug, or internal self-audit), and the time and date.
    IdenTrust:
    On November 6, 2020 at 4:37PM MT, the OCSP responder monitoring system we implemented as a result of bug 1636544 alerted of abnormal behavior. Upon review of the monitor and testing, IdenTrust engineers did not immediately identify a problem as connectivity to all IdenTrust services could be made.

On November 7, 2020 at 1:45AM MT, IdenTrust was made aware of a customer reporting an intermittent error while connecting to the OCSP. Upon investigating with the customer, we identified a degradation in service for traffic accessing the secondary data center.
2. A timeline of the actions your CA took in response. A timeline is a date-and-time-stamped sequence of all relevant events. This may include events before the incident was reported, such as when a particular requirement became applicable, or a document changed, or a bug was introduced, or an audit was done.
IdenTrust:
November 6, 2020 at 4:37 PM MT: Received alert from OCSP responder monitoring system
November 6, 2020 at 4:45 PM MT: Engineers conducted testing of services and no failures were reported
November 6, 2020 at 8:13 PM MT: Receive alert from OCSP responder monitoring system
November 6, 2020 at 9:00 PM MT: Engineers conducted additional testing of the IdenTrust Commercial Root CA A1 again and confirmed that there were no connectivity issues.
November 7, 2020 at 1:44 AM MT: Receive alert from OCSP responder monitoring system
November 7, 2020 at 1:45 AM MT: Received external notice on failed response to resolve a DNS lookup.
November 7, 2020 at 1:56 AM MT: Engineers began testing and verifying the problem
November 7, 2020 at 2:37 AM MT: Issue was escalated internally for a full review of the problem
November 7, 2020 at 3:15 AM MT: Identified the problem and developed a plan to resolve the issue.
November 7, 2020 at 3:40 AM MT: The DNS server was removed and marked as unavailable to stop impact to customers. Testing confirmed all services were available.
3. Whether your CA has stopped, or has not yet stopped, issuing certificates with the problem. A statement that you have will be considered a pledge to the community; a statement that you have not requires an explanation.
IdenTrust:
Not applicable
4. A summary of the problematic certificates. For each problem: number of certs, and the date the first and last certs with that problem were issued.
IdenTrust:
Not applicable
5. The complete certificate data for the problematic certificates. The recommended way to provide this is to ensure each certificate is logged to CT and then list the fingerprints or crt.sh IDs, either in the report or as an attached spreadsheet, with one list per distinct problem.
IdenTrust:
Not applicable
6. Explanation about how and why the mistakes were made or bugs introduced, and how they avoided detection until now.
IdenTrust:
IdenTrust maintains 4 DNS servers, and on November 7, 2020 , 1 of those 4 servers became unstable due a hardware controller error resulting in a degraded service. Hosted in this environment is a DNS server that continued to respond to queries with an empty value. This caused some validations attempts fail to reach the OCSP responder as the DNS lookup failed. This impacted up to 25% of incoming traffic. The failure prevented some virtual hosts from establishing connection to the database and synchronizing with other servers. As the issues were addressed, the hosts impacted were re-introduced to the server cluster. The hardware controller failure was related to a configuration within the virtual environment that resulted in a connection disconnect on the disk that the host resides on.
7. List of steps your CA is taking to resolve the situation and ensure such issuance will not be repeated in the future, accompanied with a timeline of when your CA expects to accomplish these things.
IdenTrust:
With this incident we have identified the items below as post-incident actions:
a. On November 7, 2020 the problematic DNS server was temporarily removed from service and from the global registrar. After making this change, the issue was resolved and the remaining DNS servers were handling external queries.
b. Engineers have worked with external vendors to address configuration changes that are needed within the virtual server environment. Required configurations were implemented on November 8th.
c. After evaluating the changes and working with vendors to confirm system readiness, the DNS records were updated on the global registrar on November 13, 2020 and made publicly available.
d. The 4th. DNS server was placed back in the cluster on November 14, 2020.

While this provides quite a bit of useful detail, I don't feel that there's a clear answer as to how this avoided detection or what steps have been taken to improve this: in particular, to improve detection.

The incident reports speaks to the remediation of the service, but not to further preventative steps that can detect and mitigate this issue sooner. For example, your OCSP systems appeared to have noticed the issue, but connecting it to this issue took what sounds like an external report. Similarly, your DNS systems happily provided services from a node in a degraded state. Can you speak to what systemic mitigations are being put in place to improve detection and prevention?

Flags: needinfo?(roots)
Flags: needinfo?(roots)
Flags: needinfo?(roots)

There was not an issue with detection of the problem. IdenTrust uses both internal and external monitoring and problem detection systems performed as expected. What was identified, was an extreme edge case scenario that triggered this problem and it took a little time to resolve. This issue was isolated to a specific load balancer that was in front of the DNS servers. The vendor has since provided specific changes to affect a more optimal failure mode.
Systematically we have adjusted the monitoring rules applied to the individual name servers for greater granularity of monitoring.

I believe this bug can be closed and plan to do so on or about next Friday, 5-March-2021.

Flags: needinfo?(bwilson)
Status: ASSIGNED → RESOLVED
Closed: 5 years ago
Flags: needinfo?(bwilson)
Resolution: --- → FIXED
Flags: needinfo?(roots)
Summary: IdenTrust Service Degradation → IdenTrust: Service Degradation
Product: NSS → CA Program
Whiteboard: [ca-compliance] → [ca-compliance] [ocsp-failure]
You need to log in before you can comment on or make changes to this bug.