Asseco DS / Certum: CP/CPS, Revocation Requests Mechanism, Certificate Problem Report, CRL and OCSP disruption
Categories
(CA Program :: CA Certificate Compliance, task)
Tracking
(Not tracked)
People
(Reporter: wtrapczynski, Assigned: wtrapczynski)
Details
(Whiteboard: [ca-compliance] [policy-failure])
Preliminary Incident Report
During scheduled network work on 2024-07-21, our infrastructure experienced a disruption that caused the unavailability of CP/CPS, Revocation Requests Mechanism, Certificate Problem Report, CRL and OCSP services. This violates the following BR rules:
-
The CA SHALL publicly disclose its Certificate Policy and/or Certification Practice Statement through an appropriate and readily accessible online means that is available on a 24x7 basis.
-
The CA SHALL maintain a continuous 24x7 ability to accept and respond to revocation requests and Certificate Problem Reports.
-
The CA SHALL maintain an online 24x7 Repository that application software can use to automatically check the current status of all unexpired Certificates issued by the CA.
The failure was resolved within about 1-2 hours, and since then all the services mentioned are working correctly.
This incident does not affect the issuance of certificates.
We will provide a full incident report no later than 2024-08-02.
Updated•2 years ago
|
| Assignee | ||
Comment 1•2 years ago
|
||
Incident Report
Summary
During scheduled network maintenance on 2024-07-21, our infrastructure experienced a disruption that caused the unavailability. The scope of network maintenance concerned the replacement of network interfaces for networks within the organisation (part one of the work) and for public networks (part two of the work). The problem occurred during the second part of the network maintenance and was related to the incorrect setting of the route in the configuration. As a result, the following services were not available from the internet.
The problem affected the following services which according to BR must be available 24x7:
-
CP/CPS
-
respond to revocation requests and Certificate Problem Report
-
CRL
-
OCSP
The failure was resolved within about 1-2 hours, and since then all the services mentioned are working correctly.
This incident does not affect the issuance of certificates.
Impact
Relying parties have not been able to download CP/CPS.
Relying parties have not been able to revoke certificate and submit a new request via Certificate Problem Report.
Relying parties have not been able to obtain certificate status via CRL and OCSP.
Timeline
All times are UTC.
2024-07-21:
-
22:00 Start of the first part of the network maintenance: replacement of network interfaces for internal networks within the organisation.
-
22:10 End of the first part of the network maintenance.
-
22:30 Confirmation that all services are working properly after the first part of the network maintenance.
-
22:35 Start of the second part of the network maintenance: replacement of network interfaces for public networks.
-
22:45 End of the second part of the network maintenance. The incorrect configuration has been implemented.
-
23:00 The monitoring system alerted us about the problem with setting up new connections from server networks to the Internet. The analysis has begun.
-
23:30 Confirmation that we are not getting full data from external monitoring servers.
-
23:50 Confirmation that our services are not accessible from the internet.
-
00:10 Start identifying the list of critical systems affected for switchover to the Secondary Data Center.
-
00:25 CRL services have been switched to the Secondary Data Center. CRL service started working properly
-
00:30 OCSP services have been switched to the Secondary Data Center. OCSP service started working properly
-
00:35 CP/CPS site have been switched to the Secondary Data Center. CP/CPS site started working properly.
-
01:00 The source of the initial problem was identified.
-
01:15 The initial problem has been solved. The route has been set correctly. All other services are working again.
-
01:30 All services have been switched back to the Primary Data Center.
-
02:00 Confirmation that all services are working properly.
Root Cause Analysis
Insufficient configuration checks
The main reason for the incident was an incorrectly set route for the new network interface in the configuration of the network device.
Significant omissions in network monitoring data
During network maintenance, a team of network administrators supervised monitoring systems and logs to detect possible problems.
After the misconfiguration, network administrators verified network traffic graphs in the monitoring system and logs from network devices. The anomalies were not immediately detected because the focus was on analysing the monitoring data of the edge device which also handles traffic from other subnets. Network traffic during the night from Certum's network was not significant enough to be noticed by its absence on the charts. The result of the log analysis also did not raise an alarm, because the traffic was still logged as usual. The only marker indicating that something was wrong was a note in the logs that the traffic was incomplete. However, it was overlooked in the initial phase.
Insufficient services monitoring
During network maintenance also a team of system administrators supervised monitoring systems and logs to detect possible problems.
The monitoring system alerted us about the failure, but it was not detailed enough, which made analysis and finding the cause very difficult. An additional factor that made analysis difficult at the initial stage was the fact that all services worked correctly from within the network. CP/CPS, Revocation Requests Mechanism, OCSP and CRL services, in addition to monitoring from within the network, are also monitored from several locations around the world. So, we were potentially able to quickly track down what wasn't working and why. Unfortunately, a network failure also blocked our access to external monitoring servers, and we were not receiving data to the central monitoring system.
CDN configuration
For some of our services we use CDN to cache content at the network edge. CDN configuration included only Primary Origin Server that could be manually switched between Primary Data Center and Secondary Data Center. In case of failure of Primary Origin Server located in Primary Data Center, the work of an administrator was required who could switch traffic to Secondary Origin Server located in Secondary Data Center. Having automated traffic switching would in this case significantly reduce the service downtime.
Lessons Learned
What went well
-
The monitoring system alerted us that we could quickly start the analysis that finally led us to identify and solve the problem.
-
The switchover procedure for services to Secondary Data Center worked correctly.
What didn't go well
-
Missing reduced network traffic in network device monitoring charts.
-
Missing erroneous entries in network device logs.
-
Although the monitoring system alerted us to a potential problem with the operation of services, there were not enough details to quickly identify the core of the problem.
-
CDN was not configured in automatic traffic switching mode from Primary Origin Server to Secondary Origin Server.
Where we got lucky
Action Items
| Action Item | Kind | Due Date |
|---|---|---|
| Implement additional monitoring controls for network devices and traffic | Detect | 2024-08-16 |
| Implement additional monitoring controls for services | Detect | 2024-08-16 |
| Changing the configuration of CDN to automate traffic switching between Primary Origins Server and Secondary Origin Server | Mitigate | 2024-08-16 |
| Add external monitoring for Certificate Problem Report | Detect | Completed 2024-07-25 |
Appendix
Details of affected certificates
This incident has not led to any misissued certificates.
Comment 2•2 years ago
|
||
We have no updates on this bug.
Comment 3•1 year ago
|
||
We have no updates on this bug.
Comment 4•1 year ago
|
||
Action Items Update
| Action Item | Kind | Due Date |
|---|---|---|
| Implement additional monitoring controls for network devices and traffic | Detect | 2024-08-14 (Completed) |
| Implement additional monitoring controls for services | Detect | 2024-08-14 (Completed) |
| Changing the configuration of CDN to automate traffic switching between Primary Origins Server and Secondary Origin Server | Mitigate | 2024-08-14 (Completed) |
| Add external monitoring for Certificate Problem Report | Detect | 2024-07-25 (Completed) |
| Assignee | ||
Comment 5•1 year ago
|
||
We have no updates on this bug.
Comment 6•1 year ago
|
||
All action items listed in the bug have now been completed, and we continue to monitor it for any questions from the community.
Comment 7•1 year ago
|
||
I will look at closing this on Friday, 30-Aug-2024.
Thanks,
Ben
Comment 8•1 year ago
|
||
As there have been no further questions or updates, could you please look into closing it at your earliest convenience?
Thanks,
Kateryna
Updated•1 year ago
|
Updated•1 year ago
|
Description
•