Kotlin Mutex withLock Deadlock: Why a Nested withLock Hangs Forever

A Mutex withLock deadlock happens when code holding a Mutex calls a function that locks it again. Two causes stack up. First, the kotlinx.coroutines Mutex is non-reentrant. The holder gets no pass. Its second lock waits like any other caller. It’s waiting on itself, so it never gets the lock. Second, the hang is silent. A suspended coroutine doesn’t block a thread. So the main thread keeps drawing frames and no ANR fires. Nothing throws. lock has no timeout either. The screen just shows a spinner forever. The usual trigger is a migration from @Synchronized. That lock is reentrant, so the nested call used to work. The fix is structural. Keep withLock only in the public entry points. Move the shared work into a private function that assumes the caller holds the lock. Then call that private function from inside the lock.

A token store that never returns its first token

The running example is an auth token cache. Many requests need a token at once. Only one refresh may run. The Mutex serializes them. validToken() locks, checks expiry and refreshes if needed. refresh() is public too, because a 401 handler calls it directly. So it locks as well. The code was a @Synchronized class before the move to suspend functions.

import kotlinx.coroutines.*
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock

class TokenStore(private val api: AuthApi) {
    private val mutex = Mutex()
    private var token: String? = null
    private var expiresAt = 0L

    suspend fun validToken(): String = mutex.withLock {
        if (token == null || System.currentTimeMillis() >= expiresAt) {
            refresh()
        }
        token!!
    }

    suspend fun refresh() = mutex.withLock {
        val fresh = api.refresh()
        token = fresh
        expiresAt = System.currentTimeMillis() + 60_000
    }
}

class AuthApi {
    suspend fun refresh(): String { delay(50); return "tok-1" }
}

fun main() = runBlocking {
    val store = TokenStore(AuthApi())
    val job = launch {
        println("requesting token")
        println("got ${store.validToken()}")
    }
    delay(2_000)
    println("after 2s: job.isActive=${job.isActive}, isCompleted=${job.isCompleted}")
    job.cancel()
    job.join()
    println("cancelled: ${job.isCancelled}")
}

Compiled with kotlinc 2.4.20 against kotlinx.coroutines 1.10.2, it prints this:

requesting token
after 2s: job.isActive=true, isCompleted=false
cancelled: true

The line got ... never appears. The fake API answers in 50 ms, yet two seconds later the job is still active. launch inside runBlocking runs on the same single thread as main. That thread still ran the delay, the print and the cancel. On Android, that thread would be the main thread. It would keep handling taps and frames while the token request never finishes.

A Mutex withLock deadlock happens because the holder waits on itself

The Mutex API docs state it directly. The mutex is non-reentrant. Calling lock suspends the caller even from the thread or coroutine that holds the lock. Walk validToken() through that rule:

  1. validToken() calls withLock. The mutex is free, so the coroutine takes it.
  2. The token is null, so it calls refresh().
  3. refresh() calls withLock on the same mutex. The mutex is locked, so the coroutine suspends and joins the wait queue.
  4. The lock can only be released when the outer withLock block ends. That block is waiting for refresh() to return. Neither side can move.

Java’s synchronized works differently. It tracks the owning thread, so the same thread can enter again. That’s why the old @Synchronized version worked. A coroutine has no fixed thread to track. It can resume on a different thread after any suspension. Making the mutex reentrant would need a stable coroutine identity instead.

The library team has declined to add one. A user filed kotlinx.coroutines issue #1686 in December 2019. It asked for a reentrant lock keyed on the current Job. Roman Elizarov of JetBrains closed it the same day. He pointed out that withContext and coroutineScope create a new Job without starting a new coroutine. So a Job isn’t a reliable identity. He called coroutines “too ephemeral” for a reentrant lock. The issue is closed but still drew comments through 2024.

The hang is silent because suspension blocks no thread

The second cause decides how long the bug survives. A thread deadlock in Java is loud. In the classic two-lock case, both threads sit in BLOCKED. A jstack dump reports a Java-level deadlock. On Android, a main thread blocked for 5 seconds on input triggers an ANR. The Android ANR docs list that input-dispatch timeout as the common foreground trigger.

A suspended coroutine triggers none of this. It’s a continuation sitting in the mutex’s wait queue. No thread is parked on it. The thread dump shows idle threads. The ANR watchdog sees a main thread that answers input on time. The lock docs describe it as suspending until the lock is acquired, with no timeout parameter. So the wait has no end.

Only cancellation ends it. The same docs say lock is cancellable. A waiting caller resumes with CancellationException. That matches the last line of the repro. In an app, viewModelScope is canceled when the ViewModel is cleared. So the bug looks like a slow network that recovers after the user leaves the screen. To see the real state, pause in the debugger. Kotlin’s coroutine debugging tutorial describes a Coroutines tab that lists each coroutine with its status. A coroutine stuck in this deadlock would show there as suspended.

Lock once at the entry point and call a private unlocked function

The fix separates two jobs. The public functions acquire the lock. A private function does the work and assumes the caller already holds the lock. The Locked suffix is only a naming habit. It warns the next reader not to call it from outside.

import kotlinx.coroutines.*
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import java.util.concurrent.atomic.AtomicInteger

class TokenStore(private val api: AuthApi) {
    private val mutex = Mutex()
    private var token: String? = null
    private var expiresAt = 0L

    suspend fun validToken(): String = mutex.withLock {
        if (token == null || System.currentTimeMillis() >= expiresAt) {
            refreshLocked()
        }
        token!!
    }

    suspend fun refresh() = mutex.withLock { refreshLocked() }

    // Caller must hold mutex.
    private suspend fun refreshLocked() {
        val fresh = api.refresh()
        token = fresh
        expiresAt = System.currentTimeMillis() + 60_000
    }
}

class AuthApi {
    val calls = AtomicInteger()
    suspend fun refresh(): String { delay(50); return "tok-${calls.incrementAndGet()}" }
}

fun main() = runBlocking {
    val api = AuthApi()
    val store = TokenStore(api)
    val tokens = List(10) { async(Dispatchers.Default) { store.validToken() } }.awaitAll()
    println("tokens=${tokens.toSet()} refreshCalls=${api.calls.get()}")
    store.refresh()
    println("after forced refresh: ${store.validToken()} refreshCalls=${api.calls.get()}")
}
tokens=[tok-1] refreshCalls=1
after forced refresh: tok-2 refreshCalls=2

Ten coroutines on Dispatchers.Default asked for a token at once. All ten got tok-1. The API was called once. The first caller refreshed while holding the lock. The other nine waited in the queue. Each then saw a fresh token and skipped the refresh. The forced refresh from the 401 path still works. It still runs under the lock.

The fix keeps one rule. Only public functions call withLock. They never call each other while holding it. If validToken() grows a new helper, the helper gets the Locked treatment too. For a different mutex guarding pagination state, see our pagination race condition post. There the bug is a missing lock.

The answer to give when the interviewer says the request never returns

This snippet fits a bug-squash or debugging round. The symptom is vague on purpose. Nothing crashes. The test just times out. Our Android debugging interview post covers how that round is scored. Here’s the explanation to give.

  1. Name the cycle. validToken() holds the mutex and calls refresh(). That function locks the same mutex. The inner call waits for a lock its own caller holds.
  2. Name the rule behind it. Mutex is non-reentrant by design. synchronized is reentrant because it tracks a thread. A coroutine has no fixed thread, so the library won’t guess an owner.
  3. Explain why nothing failed loudly. The coroutine is suspended, not blocked. No thread is stuck, so no ANR and no thread-dump deadlock. lock waits until cancellation.
  4. Give the fix. Lock only at public entry points. Move the shared body into a private function that requires the caller to hold the lock. Call that function from both entry points.
  5. Offer a test. Call validToken() on an empty cache inside withTimeout. The nested lock then fails the test in seconds instead of hanging it.

Point 3 explains why the bug reached production. It also shows you know how suspension differs from blocking.

Five fixes that hide the hang or move it somewhere else

Each of these came up against the token store. The ones with a run behind them used the same compiler and library versions as above.

  • Going back to synchronized won’t compile here. With delay inside a synchronized block, kotlinc 2.4.20 reports that the suspension point is inside a critical section. The lock is tied to a thread. A suspension can move the code to another one.
  • A java.util.concurrent.locks.ReentrantLock is reentrant per thread. The ReentrantLock docs say the lock is owned by the thread that locked it. unlock from any other thread throws IllegalMonitorStateException. A coroutine can resume on another thread after the network call. Then its unlock throws. Its lock also blocks a real thread while it waits.
  • Skipping the lock when mutex.isLocked is true fails differently. isLocked says someone holds the mutex. It doesn’t say who. A second coroutine would see true and run refreshLocked() with no lock at all. That brings back the double refresh the mutex was there to stop.
  • Wrapping the call in withTimeout only changes the symptom. The inner lock can never succeed. The token is never set, so the next call nests again. Two calls with a one-second timeout both ended in TimeoutCancellationException. The hang became a failure on every call.
  • Shipping withLock(owner = this) breaks normal callers. The lock docs call the owner an optional token for debugging. A same-owner nested lock throws IllegalStateException. Our nested run printed This mutex is already locked by the specified owner. That’s a good debug signal. But this is the same store for every caller. When two separate coroutines called refresh() at once, the second one threw IllegalStateException too. It should have waited.

A reentrant wrapper does exist. Elizarov later posted a context-based version as a gist in the same issue thread. He also told one user there that they would likely redesign. He expected a design with no reentrant lock at all. For a token store, that redesign is the split into locked and unlocked functions.

Leave a Comment