A Mutex withLock deadlock happens when code holding a Mutex calls a function that locks it again. Two causes stack up. First, the kotlinx.coroutines Mutex is non-reentrant. The holder gets no pass. Its second lock waits like any other caller. It’s waiting on itself, so it never gets the lock. Second, the hang is silent. A suspended coroutine doesn’t block a thread. So the main thread keeps drawing frames and no ANR fires. Nothing throws. lock has no timeout either. The screen just shows a spinner forever. The usual trigger is a migration from @Synchronized. That lock is reentrant, so the nested call used to work. The fix is structural. Keep withLock only in the public entry points. Move the shared work into a private function that assumes the caller holds the lock. Then call that private function from inside the lock.
A token store that never returns its first token
The running example is an auth token cache. Many requests need a token at once. Only one refresh may run. The Mutex serializes them. validToken() locks, checks expiry and refreshes if needed. refresh() is public too, because a 401 handler calls it directly. So it locks as well. The code was a @Synchronized class before the move to suspend functions.
import kotlinx.coroutines.*
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
class TokenStore(private val api: AuthApi) {
private val mutex = Mutex()
private var token: String? = null
private var expiresAt = 0L
suspend fun validToken(): String = mutex.withLock {
if (token == null || System.currentTimeMillis() >= expiresAt) {
refresh()
}
token!!
}
suspend fun refresh() = mutex.withLock {
val fresh = api.refresh()
token = fresh
expiresAt = System.currentTimeMillis() + 60_000
}
}
class AuthApi {
suspend fun refresh(): String { delay(50); return "tok-1" }
}
fun main() = runBlocking {
val store = TokenStore(AuthApi())
val job = launch {
println("requesting token")
println("got ${store.validToken()}")
}
delay(2_000)
println("after 2s: job.isActive=${job.isActive}, isCompleted=${job.isCompleted}")
job.cancel()
job.join()
println("cancelled: ${job.isCancelled}")
}
Compiled with kotlinc 2.4.20 against kotlinx.coroutines 1.10.2, it prints this:
requesting token after 2s: job.isActive=true, isCompleted=false cancelled: true
The line got ... never appears. The fake API answers in 50 ms, yet two seconds later the job is still active. launch inside runBlocking runs on the same single thread as main. That thread still ran the delay, the print and the cancel. On Android, that thread would be the main thread. It would keep handling taps and frames while the token request never finishes.
A Mutex withLock deadlock happens because the holder waits on itself
The Mutex API docs state it directly. The mutex is non-reentrant. Calling lock suspends the caller even from the thread or coroutine that holds the lock. Walk validToken() through that rule:
validToken()callswithLock. The mutex is free, so the coroutine takes it.- The token is null, so it calls
refresh(). refresh()callswithLockon the same mutex. The mutex is locked, so the coroutine suspends and joins the wait queue.- The lock can only be released when the outer
withLockblock ends. That block is waiting forrefresh()to return. Neither side can move.
Java’s synchronized works differently. It tracks the owning thread, so the same thread can enter again. That’s why the old @Synchronized version worked. A coroutine has no fixed thread to track. It can resume on a different thread after any suspension. Making the mutex reentrant would need a stable coroutine identity instead.
The library team has declined to add one. A user filed kotlinx.coroutines issue #1686 in December 2019. It asked for a reentrant lock keyed on the current Job. Roman Elizarov of JetBrains closed it the same day. He pointed out that withContext and coroutineScope create a new Job without starting a new coroutine. So a Job isn’t a reliable identity. He called coroutines “too ephemeral” for a reentrant lock. The issue is closed but still drew comments through 2024.
The hang is silent because suspension blocks no thread
The second cause decides how long the bug survives. A thread deadlock in Java is loud. In the classic two-lock case, both threads sit in BLOCKED. A jstack dump reports a Java-level deadlock. On Android, a main thread blocked for 5 seconds on input triggers an ANR. The Android ANR docs list that input-dispatch timeout as the common foreground trigger.
A suspended coroutine triggers none of this. It’s a continuation sitting in the mutex’s wait queue. No thread is parked on it. The thread dump shows idle threads. The ANR watchdog sees a main thread that answers input on time. The lock docs describe it as suspending until the lock is acquired, with no timeout parameter. So the wait has no end.
Only cancellation ends it. The same docs say lock is cancellable. A waiting caller resumes with CancellationException. That matches the last line of the repro. In an app, viewModelScope is canceled when the ViewModel is cleared. So the bug looks like a slow network that recovers after the user leaves the screen. To see the real state, pause in the debugger. Kotlin’s coroutine debugging tutorial describes a Coroutines tab that lists each coroutine with its status. A coroutine stuck in this deadlock would show there as suspended.
Lock once at the entry point and call a private unlocked function
The fix separates two jobs. The public functions acquire the lock. A private function does the work and assumes the caller already holds the lock. The Locked suffix is only a naming habit. It warns the next reader not to call it from outside.
import kotlinx.coroutines.*
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import java.util.concurrent.atomic.AtomicInteger
class TokenStore(private val api: AuthApi) {
private val mutex = Mutex()
private var token: String? = null
private var expiresAt = 0L
suspend fun validToken(): String = mutex.withLock {
if (token == null || System.currentTimeMillis() >= expiresAt) {
refreshLocked()
}
token!!
}
suspend fun refresh() = mutex.withLock { refreshLocked() }
// Caller must hold mutex.
private suspend fun refreshLocked() {
val fresh = api.refresh()
token = fresh
expiresAt = System.currentTimeMillis() + 60_000
}
}
class AuthApi {
val calls = AtomicInteger()
suspend fun refresh(): String { delay(50); return "tok-${calls.incrementAndGet()}" }
}
fun main() = runBlocking {
val api = AuthApi()
val store = TokenStore(api)
val tokens = List(10) { async(Dispatchers.Default) { store.validToken() } }.awaitAll()
println("tokens=${tokens.toSet()} refreshCalls=${api.calls.get()}")
store.refresh()
println("after forced refresh: ${store.validToken()} refreshCalls=${api.calls.get()}")
}
tokens=[tok-1] refreshCalls=1 after forced refresh: tok-2 refreshCalls=2
Ten coroutines on Dispatchers.Default asked for a token at once. All ten got tok-1. The API was called once. The first caller refreshed while holding the lock. The other nine waited in the queue. Each then saw a fresh token and skipped the refresh. The forced refresh from the 401 path still works. It still runs under the lock.
The fix keeps one rule. Only public functions call withLock. They never call each other while holding it. If validToken() grows a new helper, the helper gets the Locked treatment too. For a different mutex guarding pagination state, see our pagination race condition post. There the bug is a missing lock.
The answer to give when the interviewer says the request never returns
This snippet fits a bug-squash or debugging round. The symptom is vague on purpose. Nothing crashes. The test just times out. Our Android debugging interview post covers how that round is scored. Here’s the explanation to give.
- Name the cycle.
validToken()holds the mutex and callsrefresh(). That function locks the same mutex. The inner call waits for a lock its own caller holds. - Name the rule behind it.
Mutexis non-reentrant by design.synchronizedis reentrant because it tracks a thread. A coroutine has no fixed thread, so the library won’t guess an owner. - Explain why nothing failed loudly. The coroutine is suspended, not blocked. No thread is stuck, so no ANR and no thread-dump deadlock.
lockwaits until cancellation. - Give the fix. Lock only at public entry points. Move the shared body into a private function that requires the caller to hold the lock. Call that function from both entry points.
- Offer a test. Call
validToken()on an empty cache insidewithTimeout. The nested lock then fails the test in seconds instead of hanging it.
Point 3 explains why the bug reached production. It also shows you know how suspension differs from blocking.
Five fixes that hide the hang or move it somewhere else
Each of these came up against the token store. The ones with a run behind them used the same compiler and library versions as above.
- Going back to
synchronizedwon’t compile here. Withdelayinside asynchronizedblock,kotlinc2.4.20 reports that the suspension point is inside a critical section. The lock is tied to a thread. A suspension can move the code to another one. - A
java.util.concurrent.locks.ReentrantLockis reentrant per thread. The ReentrantLock docs say the lock is owned by the thread that locked it.unlockfrom any other thread throwsIllegalMonitorStateException. A coroutine can resume on another thread after the network call. Then itsunlockthrows. Itslockalso blocks a real thread while it waits. - Skipping the lock when
mutex.isLockedis true fails differently.isLockedsays someone holds the mutex. It doesn’t say who. A second coroutine would seetrueand runrefreshLocked()with no lock at all. That brings back the double refresh the mutex was there to stop. - Wrapping the call in
withTimeoutonly changes the symptom. The innerlockcan never succeed. The token is never set, so the next call nests again. Two calls with a one-second timeout both ended inTimeoutCancellationException. The hang became a failure on every call. - Shipping
withLock(owner = this)breaks normal callers. Thelockdocs call the owner an optional token for debugging. A same-owner nested lock throwsIllegalStateException. Our nested run printedThis mutex is already locked by the specified owner. That’s a good debug signal. Butthisis the same store for every caller. When two separate coroutines calledrefresh()at once, the second one threwIllegalStateExceptiontoo. It should have waited.
A reentrant wrapper does exist. Elizarov later posted a context-based version as a gist in the same issue thread. He also told one user there that they would likely redesign. He expected a design with no reentrant lock at all. For a token store, that redesign is the split into locked and unlocked functions.