SwissSign: OCSP responder unreachable
Categories
(CA Program :: CA Certificate Compliance, task)
Tracking
(Not tracked)
People
(Reporter: michael.guenther, Assigned: michael.guenther)
Details
(Whiteboard: [ca-compliance] [ocsp-failure])
User Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:80.0) Gecko/20100101 Firefox/80.0
Steps to reproduce:
How your CA first became aware of the problem (e.g. via a problem report submitted to your Problem Reporting Mechanism, a discussion in mozilla.dev.security.policy, a Bugzilla bug, or internal self-audit), and the time and date.
Internal monitoring detected that OCSP responder was not reachable anymore.
A timeline of the actions your CA took in response. A timeline is a date-and-time-stamped sequence of all relevant events. This may include events before the incident was reported, such as when a particular requirement became applicable, or a document changed, or a bug was introduced, or an audit was done.
20200831 08:50:45 CEST OCSP service is down
20200831 ~08:55 CEST Analysis by IT OPS started
20200831 ~13:00 CEST IT Ops identified root cause: BPG configuration we always used was no longer valid
20200831 13:05 CEST Information of our Auditors
20200831 ~13:10 CEST Starting reconfiguration of BGP configuration and propagating
20200831 14:11:13 CEST OCSP up
20200831 14:16:14 CEST OCSP down again because of hardware failure.
20200831 16:12:59 CEST OCSP up: Routing to our site seems to be restored properly
20200831 16:23:59 CEST OCSP down
20200831 ~16:30 CEST OCSP service still down
Whether your CA has stopped, or has not yet stopped, issuing certificates with the problem. A statement that you have will be considered a pledge to the community; a statement that you have not requires an explanation.
n/a
A summary of the problematic certificates. For each problem: number of certs, and the date the first and last certs with that problem were issued.
n/a
The complete certificate data for the problematic certificates. The recommended way to provide this is to ensure each certificate is logged to CT and then list the fingerprints or crt.sh IDs, either in the report or as an attached spreadsheet, with one list per distinct problem.
n/a
Explanation about how and why the mistakes were made or bugs introduced, and how they avoided detection until now.
BGP changes by providers due to the outage of this weekend (Centurylink) started invalidated our BGP peering.
These resulted in a slow outage of the SwissSign Internet connection culminating this morning.
Since the configuration has not changed in years we were looking for other root causes. After fixing the BGP configuration, routes started to propagate again. After a few minutes we had a massive hardware failure on the external firewalls.
List of steps your CA is taking to resolve the situation and ensure such issuance will not be repeated in the future, accompanied with a timeline of when your CA expects to accomplish these things.
Change BPG configuration and routing. Replace failed firewall hardware cluster
I will update this ticket tomorrow with additional information.
| Assignee | ||
Updated•5 years ago
|
Updated•5 years ago
|
| Assignee | ||
Comment 1•5 years ago
|
||
Information since yesterday
Updated timeline
20200831 18:29 CEST OCSP is up. Some minor hickups during Tuesday morning of 6 minutes.
20200901 ~08:30 CEST We were able to confirm our suspicion that we are under a massive DDOS attack (up to 22Gbit)
20200901 ~10:00 Working with our ISPs on options to defend against and preventing the attack
20200901 ~13:00 CEST rules set to prevent DDOS attack is going live with positive results.
Current state: services are available and stable. DDOS attack is still ongoing
Updated Root cause analysis
Based on the additional information from today, we discovered that the DDOS protection we had in place was not enough to cope with the load of the attacks, and that the hardware failure from yesterday is connected to the DDOS. Additional measures with our internet providers have been put in place to block the same type of attack patterns we saw during the DDOS.
| Assignee | ||
Comment 2•5 years ago
|
||
Updated timeline
20200901 20:38 CEST OCPS was down again due to DDOS during the night
20200902 was up again in the morning / but not stable – we have been facing heavy attacks during the whole day
20200902 16:17 CEST OCSP is up until today
20200903 ~04:00 CEST implementation of additional external DDOS protection by Akamai
Additional DDOS attacks (up to 100Gb) hit the Akamai protection. This had no more impact on our OCSP service.
We will monitor during the weekend and make a next update on Monday/Tuesday
| Assignee | ||
Comment 3•5 years ago
|
||
The Akamai protection is performing well in the last days. There were multiple heavy attacks but these did not have any significant impact on our systems.
If there are no questions I ask that this ticket is closed.
Comment 4•5 years ago
|
||
Are there any additional questions? If not, I propose that we close this bug on or about 14-Sept-2020.
Updated•5 years ago
|
Updated•3 years ago
|
Updated•3 years ago
|
Description
•