{"id":647,"date":"2026-10-01T14:10:00","date_gmt":"2026-10-01T12:10:00","guid":{"rendered":"https:\/\/grindloop.io\/blog\/?p=647"},"modified":"2026-09-27T19:29:50","modified_gmt":"2026-09-27T17:29:50","slug":"offline-sync-duplicates-after-retry","status":"publish","type":"post","link":"https:\/\/grindloop.ai\/blog\/offline-sync-duplicates-after-retry\/","title":{"rendered":"Why Offline Sync Duplicates Records After a Retry"},"content":{"rendered":"<p>Offline sync duplicates records because a retried <code>POST<\/code> can&#8217;t tell the server it&#8217;s a retry. The server saves the note. Then the connection drops before the response arrives. The worker sees an <code>IOException<\/code> and returns <code>Result.retry()<\/code>. WorkManager replays the same <code>POST<\/code>. The server saves a second copy. RFC 9110 defines <code>POST<\/code> as non-idempotent for exactly this reason. Most candidates reach for an idempotency key next. That&#8217;s the right tool. The second bug hides in where the key gets created. If the key is generated inside the worker or an OkHttp interceptor, every attempt carries a brand-new key. The server sees two unrelated requests and saves both. The fix is to generate the key once, when the note is written locally. Store it in the outbox row. Send that same key on every attempt. The server returns the stored result when it sees the key again.<\/p>\n\n<blockquote><p>The interviewer asks for offline note-taking in an Android app. Users write notes on a train with no signal. The notes sync later. The candidate draws a Room table, a WorkManager sync worker and a <code>POST \/notes<\/code> call. Then the follow-up lands. &#8220;A user on a bad connection reports every note showing up twice on the web app. Where do the duplicates come from?&#8221;<\/p><\/blockquote>\n\n<h2 class=\"wp-block-heading\">How offline sync duplicates a note the user saved once<\/h2>\n\n<p>This is the shape most whiteboards end up with. It follows the lazy-write strategy from the <a href=\"https:\/\/developer.android.com\/topic\/architecture\/data-layer\/offline-first\" target=\"_blank\" rel=\"noopener\">Android offline-first guide<\/a>. Write locally first, then queue the network write.<\/p>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"kotlin\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">@Entity(tableName = \"outbox\")\ndata class OutboxNote(\n    @PrimaryKey(autoGenerate = true) val localId: Long = 0,\n    val text: String,\n    val createdAt: Long\n)\n\ninterface NotesApi {\n    @POST(\"notes\")\n    suspend fun createNote(@Body body: NoteDto): NoteDto\n}\n\nclass SyncWorker(\n    context: Context,\n    params: WorkerParameters,\n    private val outbox: OutboxDao,\n    private val api: NotesApi\n) : CoroutineWorker(context, params) {\n\n    override suspend fun doWork(): Result {\n        for (note in outbox.pending()) {\n            try {\n                api.createNote(NoteDto(note.text, note.createdAt))\n                outbox.delete(note)\n            } catch (e: IOException) {\n                return Result.retry()\n            }\n        }\n        return Result.success()\n    }\n}<\/pre>\n\n\n<p>Every piece here is reasonable on its own. The note survives process death because it&#8217;s in Room. The worker only deletes a row after the server accepts it. On failure it asks WorkManager to try again. The <a href=\"https:\/\/developer.android.com\/develop\/background-work\/background-tasks\/persistent\/getting-started\/define-work\" target=\"_blank\" rel=\"noopener\">WorkManager retry docs<\/a> say the default backoff is <code>EXPONENTIAL<\/code> with a 30-second delay. The minimum is 10 seconds. The offline-first guide covers queuing, backoff and conflict resolution in detail. It doesn&#8217;t discuss what a replayed write does on the server. The duplicate comes from that replay.<\/p>\n\n<h2 class=\"wp-block-heading\">Bug 1: an IOException doesn&#8217;t mean the server did nothing<\/h2>\n\n<p>A failed request has two very different causes that look identical on the client. The request may never have reached the server. Or the server may have committed the write and the response got lost on the way back. On a train, the second case is easy to hit. The client can&#8217;t distinguish them. It only sees the exception.<\/p>\n\n<p><a href=\"https:\/\/www.rfc-editor.org\/rfc\/rfc9110.html#section-9.2.2\" target=\"_blank\" rel=\"noopener\">RFC 9110, section 9.2.2<\/a> (June 2022) is the rule behind this. Of the methods it defines, only <code>PUT<\/code>, <code>DELETE<\/code> and the safe methods are idempotent. Those can be repeated after a communication failure because the effect is the same. <code>POST<\/code> isn&#8217;t on that list. The RFC says a client &#8220;SHOULD NOT automatically retry a request with a non-idempotent method.&#8221; The exception is a client with some way to know the retry is safe. <code>Result.retry()<\/code> is an automatic retry of a <code>POST<\/code> with no such way.<\/p>\n\n<p>There&#8217;s a quieter retry underneath too. OkHttp&#8217;s <a href=\"https:\/\/github.com\/square\/okhttp\/blob\/master\/okhttp\/src\/commonJvmAndroid\/kotlin\/okhttp3\/OkHttpClient.kt\" target=\"_blank\" rel=\"noopener\"><code>retryOnConnectionFailure<\/code><\/a> is on by default. Its docs say the client &#8220;silently recovers&#8221; from stale pooled connections, among other problems. The logic lives in <a href=\"https:\/\/github.com\/square\/okhttp\/blob\/master\/okhttp\/src\/commonJvmAndroid\/kotlin\/okhttp3\/internal\/http\/RetryAndFollowUpInterceptor.kt\" target=\"_blank\" rel=\"noopener\"><code>RetryAndFollowUpInterceptor<\/code><\/a>. Among its checks, it refuses to replay a request whose body is one-shot. None of the checks look at the HTTP method. So a <code>POST<\/code> can be sent twice before your worker ever sees an exception. The same OkHttp docs suggest turning it off &#8220;when doing so is destructive.&#8221; That removes one retry layer. It does nothing about the worker&#8217;s own retry after a lost response.<\/p>\n\n<h2 class=\"wp-block-heading\">Bug 2: a fresh UUID per attempt is still a new request<\/h2>\n\n<p>The textbook fix is an <code>Idempotency-Key<\/code> header. The IETF httpapi working group&#8217;s <a href=\"https:\/\/datatracker.ietf.org\/doc\/draft-ietf-httpapi-idempotency-key-header\/\" target=\"_blank\" rel=\"noopener\">Idempotency-Key draft<\/a> describes it as a way to make <code>POST<\/code> fault-tolerant. The draft is an expired Internet-Draft. Its latest version, -07, is from October 2025. It isn&#8217;t an RFC. The design is still the common pattern. The client sends a unique value. The server stores it with the result and recognizes later retries of the same request.<\/p>\n\n<p>Under interview pressure, the key often gets added in the wrong place.<\/p>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"kotlin\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">\/\/ Looks like the fix. Isn't.\nclass IdempotencyInterceptor : Interceptor {\n    override fun intercept(chain: Interceptor.Chain): Response {\n        val request = chain.request().newBuilder()\n            .header(\"Idempotency-Key\", UUID.randomUUID().toString())\n            .build()\n        return chain.proceed(request)\n    }\n}<\/pre>\n\n\n<p>The interceptor runs once per call. Each <code>Result.retry()<\/code> leads to a new <code>doWork()<\/code> run. That run makes a new call and gets a new UUID. Generating the key at the top of <code>doWork()<\/code> has the same flaw. The server commits under key A. The response is lost. The retry arrives with key B. Key B has never been seen, so the server saves the note again. The header is present but deduplicates nothing. The draft&#8217;s own rule is that a key must not be reused for a different request. The reverse holds for retries. A retry of the same request has to carry the same key.<\/p>\n\n<h2 class=\"wp-block-heading\">The failure, reproduced in plain Kotlin<\/h2>\n\n<p>This program replaces Room, WorkManager and the network with small fakes. The fake network commits on the server and then throws on the first call, like a dropped connection. The loop stands in for a worker returning <code>Result.retry()<\/code>. It compiles and runs on Kotlin 2.4.20.<\/p>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"kotlin\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">import java.io.IOException\nimport java.util.UUID\n\ndata class OutboxRow(val localId: Long, val text: String, val idempotencyKey: String? = null)\n\nclass FakeServer {\n    val notes = mutableListOf&lt;String&gt;()\n    private val seenKeys = mutableMapOf&lt;String, Int&gt;() \/\/ key -&gt; stored note id\n\n    fun createNote(text: String, key: String?): Int {\n        if (key != null) seenKeys[key]?.let { return it } \/\/ replay: return stored result\n        notes += text\n        val id = notes.size\n        if (key != null) seenKeys[key] = id\n        return id\n    }\n}\n\n\/\/ Commits on the server, then \"loses\" the first response, like a dropped connection.\nclass FlakyNetwork(private val server: FakeServer) {\n    private var calls = 0\n    fun post(text: String, key: String?): Int {\n        val id = server.createNote(text, key)\n        if (++calls == 1) throw IOException(\"connection reset before response\")\n        return id\n    }\n}\n\n\/\/ Stand-in for a CoroutineWorker that returns Result.retry() on IOException.\nfun drain(row: OutboxRow, keyFor: (OutboxRow) -&gt; String?): FakeServer {\n    val server = FakeServer()\n    val network = FlakyNetwork(server)\n    repeat(3) { attempt -&gt;\n        try {\n            network.post(row.text, keyFor(row))\n            return server\n        } catch (e: IOException) {\n            println(\"  attempt ${attempt + 1}: $e -&gt; Result.retry()\")\n        }\n    }\n    return server\n}\n\nfun main() {\n    val row = OutboxRow(localId = 1, text = \"Standup notes\")\n\n    println(\"1: no key\")\n    println(\"  server notes: ${drain(row) { null }.notes}\")\n\n    println(\"2: key generated per attempt\")\n    println(\"  server notes: ${drain(row) { UUID.randomUUID().toString() }.notes}\")\n\n    println(\"3: key stored in the outbox row at write time\")\n    val stored = row.copy(idempotencyKey = UUID.randomUUID().toString())\n    println(\"  server notes: ${drain(stored) { it.idempotencyKey }.notes}\")\n}<\/pre>\n\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"generic\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">1: no key\n  attempt 1: java.io.IOException: connection reset before response -&gt; Result.retry()\n  server notes: [Standup notes, Standup notes]\n2: key generated per attempt\n  attempt 1: java.io.IOException: connection reset before response -&gt; Result.retry()\n  server notes: [Standup notes, Standup notes]\n3: key stored in the outbox row at write time\n  attempt 1: java.io.IOException: connection reset before response -&gt; Result.retry()\n  server notes: [Standup notes]<\/pre>\n\n\n<p>Cases 1 and 2 produce the same duplicate. Adding a header changed nothing because the key changed with the attempt. Only case 3 keeps one note. The retry carries the key the server already stored. So the server returns the saved result instead of writing again.<\/p>\n\n<h2 class=\"wp-block-heading\">The fix: mint the key when the note is written, not when it&#8217;s sent<\/h2>\n\n\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"kotlin\" data-enlighter-theme=\"\" data-enlighter-highlight=\"\" data-enlighter-linenumbers=\"\" data-enlighter-lineoffset=\"\" data-enlighter-title=\"\" data-enlighter-group=\"\">@Entity(tableName = \"outbox\")\ndata class OutboxNote(\n    @PrimaryKey(autoGenerate = true) val localId: Long = 0,\n    val text: String,\n    val createdAt: Long,\n    val idempotencyKey: String = UUID.randomUUID().toString() \/\/ set once, at insert\n)\n\ninterface NotesApi {\n    @POST(\"notes\")\n    suspend fun createNote(\n        @Header(\"Idempotency-Key\") key: String,\n        @Body body: NoteDto\n    ): NoteDto\n}\n\n\/\/ In SyncWorker.doWork()\nfor (note in outbox.pending()) {\n    try {\n        api.createNote(note.idempotencyKey, NoteDto(note.text, note.createdAt))\n        outbox.delete(note)\n    } catch (e: IOException) {\n        return Result.retry()\n    }\n}<\/pre>\n\n\n<p>The key now belongs to the note, not to the HTTP call. It&#8217;s written to Room in the same insert as the text. So it survives process death, a WorkManager reschedule and an app update. Every attempt for that row sends the same value. A client-generated note ID works the same way if the API accepts one. The server can then treat the create as an upsert on that ID.<\/p>\n\n<p>The client half only works if the server holds up its end. The server has to store each key with its result for some retention window. It returns that stored result when the key comes back. The draft also covers the in-flight case. A retry that arrives while the first request is still processing should get a <code>409<\/code>. A key reused with a different payload should get a <code>422<\/code>. In an interview, say out loud that this is a contract with the backend.<\/p>\n\n<p>Two smaller details stop the fix from leaking. Delete the outbox row only after a success response. The worker above already does that. And enqueue the sync as unique work. The offline-first guide&#8217;s sample uses <code>ExistingWorkPolicy.KEEP<\/code> so only one sync worker runs at a time. The <a href=\"https:\/\/developer.android.com\/develop\/background-work\/background-tasks\/persistent\/how-to\/manage-work\" target=\"_blank\" rel=\"noopener\">WorkManager unique-work docs<\/a> say <code>REPLACE<\/code> cancels the existing work. A worker cancelled mid-request may already have reached the server. The new worker then sends the same rows again. With a stored key, that resend is harmless. Without one, it&#8217;s another duplicate.<\/p>\n\n<h2 class=\"wp-block-heading\">What a strong system design answer covers here<\/h2>\n\n<ol>\n<li>Separate &#8220;the request failed&#8221; from &#8220;the client didn&#8217;t get a response.&#8221; Say that a lost response after a commit can happen on any flaky mobile network.<\/li>\n<li>Name <code>POST<\/code> as non-idempotent. Point out that there are two retry layers, WorkManager&#8217;s <code>Result.retry()<\/code> and OkHttp&#8217;s silent connection retry.<\/li>\n<li>Propose an idempotency key or a client-generated ID. Then say where it&#8217;s created. It&#8217;s made at local write time and persisted in the outbox row.<\/li>\n<li>Describe the server side. It stores each key with its result. It replays that result for a repeat key, within a retention window.<\/li>\n<li>Mention unique work with <code>KEEP<\/code> and deleting rows only after success. Those stop a second worker from racing the first.<\/li>\n<\/ol>\n\n<p>Scoping this in time is its own skill. Hass covers that side in <a href=\"https:\/\/grindloop.ai\/blog\/what-the-senior-android-system-design-round-actually-tests\/\">what the senior Android system design round tests<\/a>. This post is the technical depth for one of its most common follow-ups.<\/p>\n\n<p><strong>Common wrong answers:<\/strong><\/p>\n<ul>\n<li>&#8220;Add a debounce on the save button.&#8221; The user tapped once. The duplicate comes from the retry, not the UI.<\/li>\n<li>&#8220;Turn off <code>retryOnConnectionFailure<\/code>.&#8221; That removes OkHttp&#8217;s retry. The worker still retries after a lost response.<\/li>\n<li>&#8220;Don&#8217;t retry <code>POST<\/code> at all.&#8221; Then a note written offline can be lost for good. That breaks the offline-first promise the design started with.<\/li>\n<li>&#8220;Generate a UUID in an interceptor.&#8221; Each attempt gets a new key, so the server can&#8217;t match the retry to the original.<\/li>\n<li>&#8220;Deduplicate on the server by comparing note text.&#8221; Two real notes can have identical text. Identity has to come from the client.<\/li>\n<\/ul>\n\n<hr\/>\n\n<p><em>Also on this site, see <a href=\"https:\/\/grindloop.ai\/blog\/viewmodelscope-async-swallows-api-call\/\">why viewModelScope.async swallows a failed API call<\/a>. It&#8217;s another network failure that behaves differently than the code suggests. GrindLoop&#8217;s System Design track turns follow-ups like this one into timed drills, each with a reviewed answer.<\/em><\/p>\n\n<p><strong>Failed the interview? Not the next one.<\/strong><\/p>","protected":false},"excerpt":{"rendered":"<p>Offline sync duplicates records when a retried POST can&#8217;t prove it&#8217;s a retry. The two stacked bugs, a runnable Kotlin repro, the fix and the interview answer.<\/p>\n","protected":false},"author":2,"featured_media":651,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"Why Offline Sync Duplicates Records After a Retry","rank_math_description":"Offline sync duplicates records when a retried POST can't prove it's a retry. The two stacked bugs, a runnable Kotlin repro, the fix and the interview answer.","rank_math_focus_keyword":"offline sync duplicates","footnotes":""},"categories":[6,40],"tags":[15,16,12,57,89,22,21,45,90],"class_list":["post-647","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-networking","category-system-design","tag-android","tag-interview-prep","tag-kotlin","tag-networking","tag-offline-first","tag-senior-interview","tag-system-design","tag-technical-interview","tag-workmanager"],"_links":{"self":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts\/647","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/comments?post=647"}],"version-history":[{"count":2,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts\/647\/revisions"}],"predecessor-version":[{"id":650,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/posts\/647\/revisions\/650"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/media\/651"}],"wp:attachment":[{"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/media?parent=647"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/categories?post=647"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/grindloop.ai\/blog\/wp-json\/wp\/v2\/tags?post=647"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}