Why Offline Sync Duplicates Records After a Retry

Offline sync duplicates records because a retried POST can’t tell the server it’s a retry. The server saves the note. Then the connection drops before the response arrives. The worker sees an IOException and returns Result.retry(). WorkManager replays the same POST. The server saves a second copy. RFC 9110 defines POST as non-idempotent for exactly this reason. Most candidates reach for an idempotency key next. That’s the right tool. The second bug hides in where the key gets created. If the key is generated inside the worker or an OkHttp interceptor, every attempt carries a brand-new key. The server sees two unrelated requests and saves both. The fix is to generate the key once, when the note is written locally. Store it in the outbox row. Send that same key on every attempt. The server returns the stored result when it sees the key again.

The interviewer asks for offline note-taking in an Android app. Users write notes on a train with no signal. The notes sync later. The candidate draws a Room table, a WorkManager sync worker and a POST /notes call. Then the follow-up lands. “A user on a bad connection reports every note showing up twice on the web app. Where do the duplicates come from?”

How offline sync duplicates a note the user saved once

This is the shape most whiteboards end up with. It follows the lazy-write strategy from the Android offline-first guide. Write locally first, then queue the network write.

@Entity(tableName = "outbox")
data class OutboxNote(
    @PrimaryKey(autoGenerate = true) val localId: Long = 0,
    val text: String,
    val createdAt: Long
)

interface NotesApi {
    @POST("notes")
    suspend fun createNote(@Body body: NoteDto): NoteDto
}

class SyncWorker(
    context: Context,
    params: WorkerParameters,
    private val outbox: OutboxDao,
    private val api: NotesApi
) : CoroutineWorker(context, params) {

    override suspend fun doWork(): Result {
        for (note in outbox.pending()) {
            try {
                api.createNote(NoteDto(note.text, note.createdAt))
                outbox.delete(note)
            } catch (e: IOException) {
                return Result.retry()
            }
        }
        return Result.success()
    }
}

Every piece here is reasonable on its own. The note survives process death because it’s in Room. The worker only deletes a row after the server accepts it. On failure it asks WorkManager to try again. The WorkManager retry docs say the default backoff is EXPONENTIAL with a 30-second delay. The minimum is 10 seconds. The offline-first guide covers queuing, backoff and conflict resolution in detail. It doesn’t discuss what a replayed write does on the server. The duplicate comes from that replay.

Bug 1: an IOException doesn’t mean the server did nothing

A failed request has two very different causes that look identical on the client. The request may never have reached the server. Or the server may have committed the write and the response got lost on the way back. On a train, the second case is easy to hit. The client can’t distinguish them. It only sees the exception.

RFC 9110, section 9.2.2 (June 2022) is the rule behind this. Of the methods it defines, only PUT, DELETE and the safe methods are idempotent. Those can be repeated after a communication failure because the effect is the same. POST isn’t on that list. The RFC says a client “SHOULD NOT automatically retry a request with a non-idempotent method.” The exception is a client with some way to know the retry is safe. Result.retry() is an automatic retry of a POST with no such way.

There’s a quieter retry underneath too. OkHttp’s retryOnConnectionFailure is on by default. Its docs say the client “silently recovers” from stale pooled connections, among other problems. The logic lives in RetryAndFollowUpInterceptor. Among its checks, it refuses to replay a request whose body is one-shot. None of the checks look at the HTTP method. So a POST can be sent twice before your worker ever sees an exception. The same OkHttp docs suggest turning it off “when doing so is destructive.” That removes one retry layer. It does nothing about the worker’s own retry after a lost response.

Bug 2: a fresh UUID per attempt is still a new request

The textbook fix is an Idempotency-Key header. The IETF httpapi working group’s Idempotency-Key draft describes it as a way to make POST fault-tolerant. The draft is an expired Internet-Draft. Its latest version, -07, is from October 2025. It isn’t an RFC. The design is still the common pattern. The client sends a unique value. The server stores it with the result and recognizes later retries of the same request.

Under interview pressure, the key often gets added in the wrong place.

// Looks like the fix. Isn't.
class IdempotencyInterceptor : Interceptor {
    override fun intercept(chain: Interceptor.Chain): Response {
        val request = chain.request().newBuilder()
            .header("Idempotency-Key", UUID.randomUUID().toString())
            .build()
        return chain.proceed(request)
    }
}

The interceptor runs once per call. Each Result.retry() leads to a new doWork() run. That run makes a new call and gets a new UUID. Generating the key at the top of doWork() has the same flaw. The server commits under key A. The response is lost. The retry arrives with key B. Key B has never been seen, so the server saves the note again. The header is present but deduplicates nothing. The draft’s own rule is that a key must not be reused for a different request. The reverse holds for retries. A retry of the same request has to carry the same key.

The failure, reproduced in plain Kotlin

This program replaces Room, WorkManager and the network with small fakes. The fake network commits on the server and then throws on the first call, like a dropped connection. The loop stands in for a worker returning Result.retry(). It compiles and runs on Kotlin 2.4.20.

import java.io.IOException
import java.util.UUID

data class OutboxRow(val localId: Long, val text: String, val idempotencyKey: String? = null)

class FakeServer {
    val notes = mutableListOf<String>()
    private val seenKeys = mutableMapOf<String, Int>() // key -> stored note id

    fun createNote(text: String, key: String?): Int {
        if (key != null) seenKeys[key]?.let { return it } // replay: return stored result
        notes += text
        val id = notes.size
        if (key != null) seenKeys[key] = id
        return id
    }
}

// Commits on the server, then "loses" the first response, like a dropped connection.
class FlakyNetwork(private val server: FakeServer) {
    private var calls = 0
    fun post(text: String, key: String?): Int {
        val id = server.createNote(text, key)
        if (++calls == 1) throw IOException("connection reset before response")
        return id
    }
}

// Stand-in for a CoroutineWorker that returns Result.retry() on IOException.
fun drain(row: OutboxRow, keyFor: (OutboxRow) -> String?): FakeServer {
    val server = FakeServer()
    val network = FlakyNetwork(server)
    repeat(3) { attempt ->
        try {
            network.post(row.text, keyFor(row))
            return server
        } catch (e: IOException) {
            println("  attempt ${attempt + 1}: $e -> Result.retry()")
        }
    }
    return server
}

fun main() {
    val row = OutboxRow(localId = 1, text = "Standup notes")

    println("1: no key")
    println("  server notes: ${drain(row) { null }.notes}")

    println("2: key generated per attempt")
    println("  server notes: ${drain(row) { UUID.randomUUID().toString() }.notes}")

    println("3: key stored in the outbox row at write time")
    val stored = row.copy(idempotencyKey = UUID.randomUUID().toString())
    println("  server notes: ${drain(stored) { it.idempotencyKey }.notes}")
}
1: no key
  attempt 1: java.io.IOException: connection reset before response -> Result.retry()
  server notes: [Standup notes, Standup notes]
2: key generated per attempt
  attempt 1: java.io.IOException: connection reset before response -> Result.retry()
  server notes: [Standup notes, Standup notes]
3: key stored in the outbox row at write time
  attempt 1: java.io.IOException: connection reset before response -> Result.retry()
  server notes: [Standup notes]

Cases 1 and 2 produce the same duplicate. Adding a header changed nothing because the key changed with the attempt. Only case 3 keeps one note. The retry carries the key the server already stored. So the server returns the saved result instead of writing again.

The fix: mint the key when the note is written, not when it’s sent

@Entity(tableName = "outbox")
data class OutboxNote(
    @PrimaryKey(autoGenerate = true) val localId: Long = 0,
    val text: String,
    val createdAt: Long,
    val idempotencyKey: String = UUID.randomUUID().toString() // set once, at insert
)

interface NotesApi {
    @POST("notes")
    suspend fun createNote(
        @Header("Idempotency-Key") key: String,
        @Body body: NoteDto
    ): NoteDto
}

// In SyncWorker.doWork()
for (note in outbox.pending()) {
    try {
        api.createNote(note.idempotencyKey, NoteDto(note.text, note.createdAt))
        outbox.delete(note)
    } catch (e: IOException) {
        return Result.retry()
    }
}

The key now belongs to the note, not to the HTTP call. It’s written to Room in the same insert as the text. So it survives process death, a WorkManager reschedule and an app update. Every attempt for that row sends the same value. A client-generated note ID works the same way if the API accepts one. The server can then treat the create as an upsert on that ID.

The client half only works if the server holds up its end. The server has to store each key with its result for some retention window. It returns that stored result when the key comes back. The draft also covers the in-flight case. A retry that arrives while the first request is still processing should get a 409. A key reused with a different payload should get a 422. In an interview, say out loud that this is a contract with the backend.

Two smaller details stop the fix from leaking. Delete the outbox row only after a success response. The worker above already does that. And enqueue the sync as unique work. The offline-first guide’s sample uses ExistingWorkPolicy.KEEP so only one sync worker runs at a time. The WorkManager unique-work docs say REPLACE cancels the existing work. A worker cancelled mid-request may already have reached the server. The new worker then sends the same rows again. With a stored key, that resend is harmless. Without one, it’s another duplicate.

What a strong system design answer covers here

  1. Separate “the request failed” from “the client didn’t get a response.” Say that a lost response after a commit can happen on any flaky mobile network.
  2. Name POST as non-idempotent. Point out that there are two retry layers, WorkManager’s Result.retry() and OkHttp’s silent connection retry.
  3. Propose an idempotency key or a client-generated ID. Then say where it’s created. It’s made at local write time and persisted in the outbox row.
  4. Describe the server side. It stores each key with its result. It replays that result for a repeat key, within a retention window.
  5. Mention unique work with KEEP and deleting rows only after success. Those stop a second worker from racing the first.

Scoping this in time is its own skill. Hass covers that side in what the senior Android system design round tests. This post is the technical depth for one of its most common follow-ups.

Common wrong answers:

  • “Add a debounce on the save button.” The user tapped once. The duplicate comes from the retry, not the UI.
  • “Turn off retryOnConnectionFailure.” That removes OkHttp’s retry. The worker still retries after a lost response.
  • “Don’t retry POST at all.” Then a note written offline can be lost for good. That breaks the offline-first promise the design started with.
  • “Generate a UUID in an interceptor.” Each attempt gets a new key, so the server can’t match the retry to the original.
  • “Deduplicate on the server by comparing note text.” Two real notes can have identical text. Identity has to come from the client.

Also on this site, see why viewModelScope.async swallows a failed API call. It’s another network failure that behaves differently than the code suggests. GrindLoop’s System Design track turns follow-ups like this one into timed drills, each with a reviewed answer.

Failed the interview? Not the next one.

Leave a Comment