Closed Bug 2033288 Opened 5 months ago Closed 5 months ago

[regression] Unable to connect to Livekit rooms

Categories

(NSS :: Libraries, defect)

defect

Tracking

(Not tracked)

RESOLVED DUPLICATE of bug 2033783

People

(Reporter: me, Unassigned)

Details

(Keywords: regression)

Since somewhere around last week, I'm unable to join Livekit WebRTC rooms on Firefox Nightly (via LaSuite Meet). mozregression is not entirely consistent somehow, but the few tries that agreed all landed on https://phabricator.services.mozilla.com/D293987. I am not sure where to get more useful information, the browser console just shows a bunch of "sending candidate", eventually followed by a timeout :(

The affected instance is at https://meet.0upti.me/, feel free to use https://meet.0upti.me/tst-test-tst as a test room.

Verified this by taking libnss3 from the last nightly before https://hg.mozilla.org/integration/autoland/rev/cab7490156a4, dropping it into my current Nightly, and it works perfectly.

Keywords: regression
Regressed by: 2026577

:jschanck, since you are the author of the regressor, bug 2026577, could you take a look? Also, could you set the severity field?

For more information, please visit BugBot documentation.

Flags: needinfo?(jschanck)

I can not reproduce.
Ilya (as I wrote via the chat) do you mind sharing some logs or more information? (in the chat it's also ok :)

Thanks!

Flags: needinfo?(jschanck) → needinfo?(me)

OK so some collective headbashing later, here's what we have:

  • it seems like the issue specifically happens when establishing the RTC connection from Livekit to the machine
  • Livekit debug log says: WARN livekit.transport rtc/transport.go:924 error reading data channel {"room": "aa380d38-8b07-42d8-8603-7097647b136b", "roomID": "RMwmLMH7Trt2AP", "participant": "3c74dd42-0e29-4cd3-9367-106c6bdc33bc", "participantID": "PA2MfnnWgwfBB3", "remote": false, "transport": "PUBLISHER", "label": "_lossy", "error": "dtls timeout: read/write timeout: context deadline exceeded"}, after which the websocket connection from the client is closed and the user is yeeted from the room
  • Livekit debug log also mentions the correct remote addresses for the client, so it's talking to at least roughly the right thing
  • sometimes, a peer to peer connection with another user in the room is established correctly, so things appear to work for a bit, and then the whole thing fails because Livekit gives up on connecting, making it yeet the user from the room; this seems like a race (unconfirmed)
  • downgrading NSS to 3.122 from the previous Firefox Nightly makes things work consistently every time with no reconnects
  • I was unable to find the right combination of MOZ_LOG flags to get the client to log anything remotely interesting
  • I was also unable to figure out how to build NSS standalone in a way that it can be dropped into an existing Firefox, which would have allowed me to bisect NSS (working on that)
  • for future reference, the exact build of NSS 3.122 I'm using came from https://archive.mozilla.org/pub/firefox/nightly/2026/04/2026-04-13-22-08-54-mozilla-central/
Flags: needinfo?(me)

Actually, more details: this happens across multiple machines, on multiple networks, Windows and Linux, IPv4 and IPv6.

I don't think the NSS uplift is the regressor. I reset the NSS directory in my worktree to just before the 3.123 beta2 uplift and rebuilt (a debug build):

git reset 76e27fdaa88b45ffb740b6bbd21bc016cee3db25^ -- security/nss/

I then attempted to join https://meet.0upti.me/tst-test-tst. After 20 seconds or so I hit an assertion failure

[522221] Assertion failure: readyState == WebSocket::CLOSING (Received message while CONNECTING or CLOSED), at ./../../../dom/websocket/WebSocket.cpp:752

The browser console shows:

sending ice candidate 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm", candidate: RTCIceCandidate }
index-D_VN5VfF.js:57:54148
sending ice candidate 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm", candidate: RTCIceCandidate }
index-D_VN5VfF.js:57:54148
sending ice candidate 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm", candidate: RTCIceCandidate }
index-D_VN5VfF.js:57:54148
sending ice candidate 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm", candidate: RTCIceCandidate }
index-D_VN5VfF.js:57:54148
sending ice candidate 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm", candidate: RTCIceCandidate }
index-D_VN5VfF.js:57:54148
The resource at “https://meet.0upti.me/assets/material-icons-outlined-latin-400-normal-DZhiGvEA.woff2” preloaded with link preload was not used within a few seconds. Make sure all attributes of the preload tag are set correctly. tst-test-tst
The resource at “https://meet.0upti.me/assets/material-symbols-outlined-latin-wght-normal-DKRrZ6AO.woff2” preloaded with link preload was not used within a few seconds. Make sure all attributes of the preload tag are set correctly. tst-test-tst
pc state change: from CONNECTING to FAILED 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm" }
index-D_VN5VfF.js:59:12597
primary PC state changed 3 
Object { room: "tst-test-tst", roomID: "RM_QRaKYnTBXyEq", participant: "dbf03f3c-ef63-4477-b7c5-22ce62a97d81", pID: "PA_n8RCqvDPAVSm" }
index-D_VN5VfF.js:59:68733
WebRTC: ICE failed, see about:webrtc for more details

I'm starting to wonder if there is some race and the NSS upgrade just wiggles the timings in the right way...

Note the official lasuite Meet demo also suffers from this. Both use the Livekit embedded TURN server.
The official Livekit demo does not however. It uses Livekit Cloud with per-region TURN servers.

I am running turn-rs here, not Livekit's embedded TURN server, though I can reproduce the issue with both.

Sorry, bad assumption. If it's about DTLS I guess it's not about TURN anyway.

I have gathered some pernosco recordings.

  • Here's a working case. It's working because I commented out the lines with the DTLS 1.2 cap added in bug 2026577. dtls_HandleHelloVerifyRequest is called once and would have hit the 1.2 cap had it been in place.
  • Here's a non-working case with the DTLS 1.2 cap in place. dtls_HandleHelloVerifyRequest is not even called.

My hunch is that there's a separate issue gating the NSS regression, making debugging tricky. Anna pointed out in the non-working case that there's a failure returned from ssl3_HandleFinished due to authCertificatePending, which could perhaps point to something.

See Also: → 2033783

This seems like an NSS issue that is testable with webrtc.

Assignee: nobody → nobody
Component: WebRTC → Libraries
Product: Core → NSS

(In reply to Andreas Pehrson [:pehrsons] from comment #10)

My hunch is that there's a separate issue gating the NSS regression, making debugging tricky. Anna pointed out in the non-working case that there's a failure returned from ssl3_HandleFinished due to authCertificatePending, which could perhaps point to something.

The non-working case here is separate indeed. AIUI it happens when receiving a ClientConfiguration from the livekit(?) websocket(?) server which has forceRelay ENABLED. The js client then sets RTCIceTransportPolicy "relay" on the peer connection. For some reason all the resulting candidate pairs fail to connect. Ilya, your server was unavailable for a period today. I think I had some luck capturing the success recording right after that (and mozregression was successful in that window too), as once relay started getting forced again I haven't gotten out of that mode. Relays not working seems orthogonal -- I hit that issue as far back as I tried (I think 138). For the case where relay is not forced, it seems bug 2033783 has a working fix in the pipe. Marking as a dupe.

Status: UNCONFIRMED → RESOLVED
Closed: 5 months ago
Duplicate of bug: 2033783
Resolution: --- → DUPLICATE

For completion I took a look at livekit. Server side logic says to fallback to TURN/TLS if TCP is not supported (which in turn is a response to UDP failing) and that's what triggers forceRelay.

Thanks everyone, I'll have to poke my TURN servers tomorrow, but that is indeed a separate issue.

No longer regressed by: 2026577
See Also: 2033783 →
You need to log in before you can comment on or make changes to this bug.