SurdaticsSURDATICS
All posts

Offline field research, and the sync problem nobody mentions

Surdatics Research

Field research happens where connectivity is not. That is the whole reason ODK Collect exists: fill forms offline, sync when you can.

The collection half is solved. The half that catches people out is the sync.

Resending is correct behaviour

ODK Collect re-sends a submission whenever it is not certain the server received it. That is not a fault — it is the right design for a tool used on an unreliable connection. Losing a completed interview is far worse than sending it twice.

But it means the server has to be able to recognise a resend. If every delivery creates a new record, one interview conducted once becomes three records, and on a platform that pays per response, three payments.

The fix is to use the identifier ODK already provides. Every filled form carries a stable instanceID in its metadata. Store it, and a resend becomes recognisable rather than a second interview.

What this looks like in practice

  • A submission is matched on its ODK instance ID, not on when it arrived
  • A resend updates the record it already created rather than adding one
  • Payment is tied to the interview, not to the delivery

None of this is visible to a fieldworker, which is the point. They fill the form, they sync when they get signal, and the count at the other end matches the number of people they actually spoke to.

Why it matters beyond the count

A research dataset is only as good as its denominator. If a study reports 400 responses and 60 of them are duplicate deliveries of the same interview, every percentage in the analysis is wrong — and wrong in a way that is very hard to spot afterwards, because the duplicates are real data, just counted twice.