Summary
Worker process crashes with a native FFI panic when two jobs' ctx.connect() calls happen within ~1 second of each other in the same worker. This reliably kills every active job on that worker, including ones that were already connected and running fine.
Tested with a non-conversational, JobContext-only entrypoint (no AgentSession, no voice pipeline) — a "programmatic participant" per the docs pattern in Job lifecycle. Repro does not require any real media/webcam — just two lk dispatch create calls to two different rooms a second apart.
Environment
livekit-agents 1.8.3
livekit (rtc/ffi) 1.1.18
- OS: Windows 11
- Python 3.13.3
- Reproduced with both
lk agent dev and lk agent start (i.e. not dev-mode-specific)
Minimal repro entrypoint
from livekit.agents import AgentServer, AutoSubscribe, JobContext, cli
server = AgentServer()
@server.rtc_session(agent_name="my-agent")
async def my_agent(ctx: JobContext) -> None:
await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_ALL)
if __name__ == "__main__":
cli.run_app(server)
Repro steps
lk agent dev (or lk agent start) to register the worker.
- Fire two dispatches to two different, brand-new rooms within ~1 second of each other:
lk dispatch create --room concurrency-test-a --agent-name my-agent &
lk dispatch create --room concurrency-test-b --agent-name my-agent &
wait
- The first job connects fine. ~10-20s later, the second job's connection attempt panics the whole worker process.
Actual behavior
WARNING livekit.agents - The room connection was not established within 10 seconds after calling job_entry. ... {"room": "concurrency-test-a"}
ERROR livekit - livekit_ffi::server::room:261:livekit_ffi::server::room - timed out waiting for ReadyForRoomEventRequest after ConnectCallback (room_handle=8)
FFI Panic: invalid request: timed out waiting for ReadyForRoomEventRequest after ConnectCallback (room_handle=8)
[exited with code 15]
The entire worker process exits — this also takes down the other, already-healthy job that was connected to concurrency-test-a.
Reproduced 3/3 times: twice under lk agent dev, once under lk agent start (production mode). Confirmed not specific to our own agent code, since it reproduces with the near-bare entrypoint above and no real participants/media.
Attempted workaround (partial, still broken)
We tried serializing ctx.connect() calls to avoid the race:
asyncio.Lock() at module scope — fails immediately with RuntimeError: <asyncio.locks.Lock object ...> is bound to a different event loop, because each job clearly runs in a separate OS process (confirmed via tasklist/pid), so an in-memory lock can't coordinate across them.
- Cross-process file lock (
filelock package) — this avoids the panic, but introduces a new failure: the job holding the lock's ctx.connect() call never returns (no exception, no log line at all after Lock ... acquired), even though the participant shows as ACTIVE in lk room participants list. The lock is then held forever, and every subsequent job spins forever retrying acquire() (never timing out, never erroring).
So today there's no way to run multiple concurrent jobs on one worker without either a process-killing panic or a silent permanent hang.
Why this matters
We're building an exam-proctoring agent (candidate video joins a room, agent watches for face-count/etc.) where multiple concurrent exam sessions per worker is the normal case, not an edge case. This bug currently makes that entirely unworkable — one candidate joining while another's session is active can silently take down every other active exam session on the same worker.
Happy to provide the full agent code, logs, or a minimal repo reproducing this if useful.
Summary
Worker process crashes with a native FFI panic when two jobs'
ctx.connect()calls happen within ~1 second of each other in the same worker. This reliably kills every active job on that worker, including ones that were already connected and running fine.Tested with a non-conversational, JobContext-only entrypoint (no
AgentSession, no voice pipeline) — a "programmatic participant" per the docs pattern in Job lifecycle. Repro does not require any real media/webcam — just twolk dispatch createcalls to two different rooms a second apart.Environment
livekit-agents1.8.3livekit(rtc/ffi) 1.1.18lk agent devandlk agent start(i.e. not dev-mode-specific)Minimal repro entrypoint
Repro steps
lk agent dev(orlk agent start) to register the worker.Actual behavior
The entire worker process exits — this also takes down the other, already-healthy job that was connected to
concurrency-test-a.Reproduced 3/3 times: twice under
lk agent dev, once underlk agent start(production mode). Confirmed not specific to our own agent code, since it reproduces with the near-bare entrypoint above and no real participants/media.Attempted workaround (partial, still broken)
We tried serializing
ctx.connect()calls to avoid the race:asyncio.Lock()at module scope — fails immediately withRuntimeError: <asyncio.locks.Lock object ...> is bound to a different event loop, because each job clearly runs in a separate OS process (confirmed viatasklist/pid), so an in-memory lock can't coordinate across them.filelockpackage) — this avoids the panic, but introduces a new failure: the job holding the lock'sctx.connect()call never returns (no exception, no log line at all afterLock ... acquired), even though the participant shows asACTIVEinlk room participants list. The lock is then held forever, and every subsequent job spins forever retryingacquire()(never timing out, never erroring).So today there's no way to run multiple concurrent jobs on one worker without either a process-killing panic or a silent permanent hang.
Why this matters
We're building an exam-proctoring agent (candidate video joins a room, agent watches for face-count/etc.) where multiple concurrent exam sessions per worker is the normal case, not an edge case. This bug currently makes that entirely unworkable — one candidate joining while another's session is active can silently take down every other active exam session on the same worker.
Happy to provide the full agent code, logs, or a minimal repo reproducing this if useful.