{"id":787,"date":"2026-10-08T13:50:00","date_gmt":"2026-10-08T11:50:00","guid":{"rendered":"https:\/\/grindloop.io\/blog\/?p=787"},"modified":"2026-10-07T22:33:59","modified_gmt":"2026-10-07T20:33:59","slug":"kotlin-mutex-withlock-deadlock","status":"publish","type":"post","link":"https:\/\/grindloop.ai\/blog\/kotlin-mutex-withlock-deadlock\/","title":{"rendered":"Kotlin Mutex withLock Deadlock: Why a Nested withLock Hangs Forever"},"content":{"rendered":"<p>A Mutex withLock deadlock happens when code holding a <code>Mutex<\/code> calls a function that locks it again. Two causes stack up. First, the kotlinx.coroutines <code>Mutex<\/code> is non-reentrant. The holder gets no pass. Its second <code>lock<\/code> waits like any other caller. It&#8217;s waiting on itself, so it never gets the lock. Second, the hang is silent. A suspended coroutine doesn&#8217;t block a thread. So the main thread keeps drawing frames and no ANR fires. Nothing throws. <code>lock<\/code> has no timeout either. The screen just shows a spinner forever. The usual trigger is a migration from <code>@Synchronized<\/code>. That lock is reentrant, so the nested call used to work. The fix is structural. Keep <code>withLock<\/code> only in the public entry points. Move the shared work into a private function that assumes the caller holds the lock. Then call that private function from inside the lock.<\/p>\n\n<h2>A token store that never returns its first token<\/h2>\n\n<p>The running example is an auth token cache. Many requests need a token at once. Only one refresh may run. The <code>Mutex<\/code> serializes them. <code>validToken()<\/code> locks, checks expiry and refreshes if needed. <code>refresh()<\/code> is public too, because a 401 handler calls it directly. So it locks as well. The code was a <code>@Synchronized<\/code> class before the move to suspend functions.<\/p>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"kotlin\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">import kotlinx.coroutines.*\nimport kotlinx.coroutines.sync.Mutex\nimport kotlinx.coroutines.sync.withLock\n\nclass TokenStore(private val api: AuthApi) {\n    private val mutex = Mutex()\n    private var token: String? = null\n    private var expiresAt = 0L\n\n    suspend fun validToken(): String = mutex.withLock {\n        if (token == null || System.currentTimeMillis() &gt;= expiresAt) {\n            refresh()\n        }\n        token!!\n    }\n\n    suspend fun refresh() = mutex.withLock {\n        val fresh = api.refresh()\n        token = fresh\n        expiresAt = System.currentTimeMillis() + 60_000\n    }\n}\n\nclass AuthApi {\n    suspend fun refresh(): String { delay(50); return \"tok-1\" }\n}\n\nfun main() = runBlocking {\n    val store = TokenStore(AuthApi())\n    val job = launch {\n        println(\"requesting token\")\n        println(\"got ${store.validToken()}\")\n    }\n    delay(2_000)\n    println(\"after 2s: job.isActive=${job.isActive}, isCompleted=${job.isCompleted}\")\n    job.cancel()\n    job.join()\n    println(\"cancelled: ${job.isCancelled}\")\n}<\/pre>\n\n\n<p>Compiled with <code>kotlinc<\/code> 2.4.20 against kotlinx.coroutines 1.10.2, it prints this:<\/p>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"generic\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">requesting token\nafter 2s: job.isActive=true, isCompleted=false\ncancelled: true<\/pre>\n\n\n<p>The line <code>got ...<\/code> never appears. The fake API answers in 50 ms, yet two seconds later the job is still active. <code>launch<\/code> inside <code>runBlocking<\/code> runs on the same single thread as <code>main<\/code>. That thread still ran the <code>delay<\/code>, the print and the cancel. On Android, that thread would be the main thread. It would keep handling taps and frames while the token request never finishes.<\/p>\n\n<h2>A Mutex withLock deadlock happens because the holder waits on itself<\/h2>\n\n<p>The <a href=\"https:\/\/kotlinlang.org\/api\/kotlinx.coroutines\/kotlinx-coroutines-core\/kotlinx.coroutines.sync\/-mutex\/\" target=\"_blank\" rel=\"noopener\">Mutex API docs<\/a> state it directly. The mutex is non-reentrant. Calling <code>lock<\/code> suspends the caller even from the thread or coroutine that holds the lock. Walk <code>validToken()<\/code> through that rule:<\/p>\n\n<ol>\n<li><code>validToken()<\/code> calls <code>withLock<\/code>. The mutex is free, so the coroutine takes it.<\/li>\n<li>The token is null, so it calls <code>refresh()<\/code>.<\/li>\n<li><code>refresh()<\/code> calls <code>withLock<\/code> on the same mutex. The mutex is locked, so the coroutine suspends and joins the wait queue.<\/li>\n<li>The lock can only be released when the outer <code>withLock<\/code> block ends. That block is waiting for <code>refresh()<\/code> to return. Neither side can move.<\/li>\n<\/ol>\n\n<p>Java&#8217;s <code>synchronized<\/code> works differently. It tracks the owning thread, so the same thread can enter again. That&#8217;s why the old <code>@Synchronized<\/code> version worked. A coroutine has no fixed thread to track. It can resume on a different thread after any suspension. Making the mutex reentrant would need a stable coroutine identity instead.<\/p>\n\n<p>The library team has declined to add one. A user filed <a href=\"https:\/\/github.com\/Kotlin\/kotlinx.coroutines\/issues\/1686\" target=\"_blank\" rel=\"noopener\">kotlinx.coroutines issue #1686<\/a> in December 2019. It asked for a reentrant lock keyed on the current <code>Job<\/code>. Roman Elizarov of JetBrains closed it the same day. He pointed out that <code>withContext<\/code> and <code>coroutineScope<\/code> create a new <code>Job<\/code> without starting a new coroutine. So a <code>Job<\/code> isn&#8217;t a reliable identity. He called coroutines &#8220;too ephemeral&#8221; for a reentrant lock. The issue is closed but still drew comments through 2024.<\/p>\n\n<h2>The hang is silent because suspension blocks no thread<\/h2>\n\n<p>The second cause decides how long the bug survives. A thread deadlock in Java is loud. In the classic two-lock case, both threads sit in <code>BLOCKED<\/code>. A <code>jstack<\/code> dump reports a Java-level deadlock. On Android, a main thread blocked for 5 seconds on input triggers an ANR. The <a href=\"https:\/\/developer.android.com\/topic\/performance\/vitals\/anr\" target=\"_blank\" rel=\"noopener\">Android ANR docs<\/a> list that input-dispatch timeout as the common foreground trigger.<\/p>\n\n<p>A suspended coroutine triggers none of this. It&#8217;s a continuation sitting in the mutex&#8217;s wait queue. No thread is parked on it. The thread dump shows idle threads. The ANR watchdog sees a main thread that answers input on time. The <a href=\"https:\/\/kotlinlang.org\/api\/kotlinx.coroutines\/kotlinx-coroutines-core\/kotlinx.coroutines.sync\/-mutex\/lock.html\" target=\"_blank\" rel=\"noopener\"><code>lock<\/code> docs<\/a> describe it as suspending until the lock is acquired, with no timeout parameter. So the wait has no end.<\/p>\n\n<p>Only cancellation ends it. The same docs say <code>lock<\/code> is cancellable. A waiting caller resumes with <code>CancellationException<\/code>. That matches the last line of the repro. In an app, <code>viewModelScope<\/code> is canceled when the ViewModel is cleared. So the bug looks like a slow network that recovers after the user leaves the screen. To see the real state, pause in the debugger. Kotlin&#8217;s <a href=\"https:\/\/kotlinlang.org\/docs\/debug-coroutines-with-idea.html\" target=\"_blank\" rel=\"noopener\">coroutine debugging tutorial<\/a> describes a Coroutines tab that lists each coroutine with its status. A coroutine stuck in this deadlock would show there as suspended.<\/p>\n\n<h2>Lock once at the entry point and call a private unlocked function<\/h2>\n\n<p>The fix separates two jobs. The public functions acquire the lock. A private function does the work and assumes the caller already holds the lock. The <code>Locked<\/code> suffix is only a naming habit. It warns the next reader not to call it from outside.<\/p>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"kotlin\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">import kotlinx.coroutines.*\nimport kotlinx.coroutines.sync.Mutex\nimport kotlinx.coroutines.sync.withLock\nimport java.util.concurrent.atomic.AtomicInteger\n\nclass TokenStore(private val api: AuthApi) {\n    private val mutex = Mutex()\n    private var token: String? = null\n    private var expiresAt = 0L\n\n    suspend fun validToken(): String = mutex.withLock {\n        if (token == null || System.currentTimeMillis() &gt;= expiresAt) {\n            refreshLocked()\n        }\n        token!!\n    }\n\n    suspend fun refresh() = mutex.withLock { refreshLocked() }\n\n    \/\/ Caller must hold mutex.\n    private suspend fun refreshLocked() {\n        val fresh = api.refresh()\n        token = fresh\n        expiresAt = System.currentTimeMillis() + 60_000\n    }\n}\n\nclass AuthApi {\n    val calls = AtomicInteger()\n    suspend fun refresh(): String { delay(50); return \"tok-${calls.incrementAndGet()}\" }\n}\n\nfun main() = runBlocking {\n    val api = AuthApi()\n    val store = TokenStore(api)\n    val tokens = List(10) { async(Dispatchers.Default) { store.validToken() } }.awaitAll()\n    println(\"tokens=${tokens.toSet()} refreshCalls=${api.calls.get()}\")\n    store.refresh()\n    println(\"after forced refresh: ${store.validToken()} refreshCalls=${api.calls.get()}\")\n}<\/pre>\n\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"generic\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">tokens=[tok-1] refreshCalls=1\nafter forced refresh: tok-2 refreshCalls=2<\/pre>\n\n\n<p>Ten coroutines on <code>Dispatchers.Default<\/code> asked for a token at once. All ten got <code>tok-1<\/code>. The API was called once. The first caller refreshed while holding the lock. The other nine waited in the queue. Each then saw a fresh token and skipped the refresh. The forced refresh from the 401 path still works. It still runs under the lock.<\/p>\n\n<p>The fix keeps one rule. Only public functions call <code>withLock<\/code>. They never call each other while holding it. If <code>validToken()<\/code> grows a new helper, the helper gets the <code>Locked<\/code> treatment too. For a different mutex guarding pagination state, see our <a href=\"https:\/\/grindloop.ai\/blog\/android-pagination-race-condition\/\">pagination race condition post<\/a>. There the bug is a missing lock.<\/p>\n\n<h2>The answer to give when the interviewer says the request never returns<\/h2>\n\n<p>This snippet fits a bug-squash or debugging round. The symptom is vague on purpose. Nothing crashes. The test just times out. Our <a href=\"https:\/\/grindloop.ai\/blog\/android-debugging-interview-second-bug\/\">Android debugging interview post<\/a> covers how that round is scored. Here&#8217;s the explanation to give.<\/p>\n\n<ol>\n<li>Name the cycle. <code>validToken()<\/code> holds the mutex and calls <code>refresh()<\/code>. That function locks the same mutex. The inner call waits for a lock its own caller holds.<\/li>\n<li>Name the rule behind it. <code>Mutex<\/code> is non-reentrant by design. <code>synchronized<\/code> is reentrant because it tracks a thread. A coroutine has no fixed thread, so the library won&#8217;t guess an owner.<\/li>\n<li>Explain why nothing failed loudly. The coroutine is suspended, not blocked. No thread is stuck, so no ANR and no thread-dump deadlock. <code>lock<\/code> waits until cancellation.<\/li>\n<li>Give the fix. Lock only at public entry points. Move the shared body into a private function that requires the caller to hold the lock. Call that function from both entry points.<\/li>\n<li>Offer a test. Call <code>validToken()<\/code> on an empty cache inside <code>withTimeout<\/code>. The nested lock then fails the test in seconds instead of hanging it.<\/li>\n<\/ol>\n\n<p>Point 3 explains why the bug reached production. It also shows you know how suspension differs from blocking.<\/p>\n\n<h2>Five fixes that hide the hang or move it somewhere else<\/h2>\n\n<p>Each of these came up against the token store. The ones with a run behind them used the same compiler and library versions as above.<\/p>\n\n<ul>\n<li>Going back to <code>synchronized<\/code> won&#8217;t compile here. With <code>delay<\/code> inside a <code>synchronized<\/code> block, <code>kotlinc<\/code> 2.4.20 reports that the suspension point is inside a critical section. The lock is tied to a thread. A suspension can move the code to another one.<\/li>\n<li>A <code>java.util.concurrent.locks.ReentrantLock<\/code> is reentrant per thread. The <a href=\"https:\/\/docs.oracle.com\/en\/java\/javase\/21\/docs\/api\/java.base\/java\/util\/concurrent\/locks\/ReentrantLock.html\" target=\"_blank\" rel=\"noopener\">ReentrantLock docs<\/a> say the lock is owned by the thread that locked it. <code>unlock<\/code> from any other thread throws <code>IllegalMonitorStateException<\/code>. A coroutine can resume on another thread after the network call. Then its <code>unlock<\/code> throws. Its <code>lock<\/code> also blocks a real thread while it waits.<\/li>\n<li>Skipping the lock when <code>mutex.isLocked<\/code> is true fails differently. <code>isLocked<\/code> says someone holds the mutex. It doesn&#8217;t say who. A second coroutine would see <code>true<\/code> and run <code>refreshLocked()<\/code> with no lock at all. That brings back the double refresh the mutex was there to stop.<\/li>\n<li>Wrapping the call in <code>withTimeout<\/code> only changes the symptom. The inner <code>lock<\/code> can never succeed. The token is never set, so the next call nests again. Two calls with a one-second timeout both ended in <code>TimeoutCancellationException<\/code>. The hang became a failure on every call.<\/li>\n<li>Shipping <code>withLock(owner = this)<\/code> breaks normal callers. The <a href=\"https:\/\/kotlinlang.org\/api\/kotlinx.coroutines\/kotlinx-coroutines-core\/kotlinx.coroutines.sync\/-mutex\/lock.html\" target=\"_blank\" rel=\"noopener\"><code>lock<\/code> docs<\/a> call the owner an optional token for debugging. A same-owner nested lock throws <code>IllegalStateException<\/code>. Our nested run printed <code>This mutex is already locked by the specified owner<\/code>. That&#8217;s a good debug signal. But <code>this<\/code> is the same store for every caller. When two separate coroutines called <code>refresh()<\/code> at once, the second one threw <code>IllegalStateException<\/code> too. It should have waited.<\/li>\n<\/ul>\n\n<p>A reentrant wrapper does exist. Elizarov later posted <a href=\"https:\/\/gist.github.com\/elizarov\/9a48b9709ffd508909d34fab6786acfe\" target=\"_blank\" rel=\"noopener\">a context-based version as a gist<\/a> in the same issue thread. He also told one user there that they would likely redesign. He expected a design with no reentrant lock at all. For a token store, that redesign is the split into locked and unlocked functions.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A Mutex withLock deadlock comes from a non-reentrant Mutex locked twice, and it hangs with no ANR or exception. A runnable repro, the fix and wrong answers.<\/p>\n","protected":false},"author":2,"featured_media":789,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"Kotlin Mutex withLock Deadlock: Why a Nested withLock Hangs Forever","rank_math_description":"A Mutex withLock deadlock comes from a non-reentrant Mutex locked twice, and it hangs with no ANR or exception. A runnable repro, the fix and wrong answers.","rank_math_focus_keyword":"mutex withlock deadlock","footnotes":""},"categories":[9,3],"tags":[15,70,13,16,12,45],"class_list":["post-787","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-bug-squash","category-coroutines","tag-android","tag-concurrency","tag-coroutines","tag-interview-prep","tag-kotlin","tag-technical-interview"],"_links":{"self":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts\/787","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/comments?post=787"}],"version-history":[{"count":1,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts\/787\/revisions"}],"predecessor-version":[{"id":788,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts\/787\/revisions\/788"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/media\/789"}],"wp:attachment":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/media?parent=787"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/categories?post=787"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/tags?post=787"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}