Menu

Country data freshness: sources, fetch dates and update strategies

Country data freshness depends on three things that are usually recorded as one. Separate the source, the fetch date and the scope to know what a value means.

Published

  • data quality
  • maintenance
  • country data

Country data freshness is normally reported as a single number: the last time the dataset was updated. That number is almost always misleading, because a dataset is not refreshed in one motion.

A typical country table is assembled from several sources, fetched at different times, covering different scopes. One entry may have been corrected last month while a neighbouring entry has been untouched for years, and a single date at the top of the file cannot express that difference.

What makes country data go stale?

Three forces, and they operate on different clocks.

The first is the world changing. Names change, currencies change, administrative divisions are reorganised, and codes are added, deprecated or reserved. None of these events announces itself to a system that is not watching for it.

The second is the source changing under you. Even a source that remains accurate can revise its own presentation, retire a field, or begin including entities it previously omitted, so a refresh can import a change nobody requested.

The third is scope drift. A rule that was written for the entities a dataset contained at the time quietly becomes wrong when that set grows, and the error appears as an exception rather than as a failure.

Only the second of the three is visible in a refresh log. The other two are the reason a dataset can be punctually refreshed and still be wrong.

Record the source, the fetch date and the scope separately

Three attributes explain a value’s provenance, and collapsing them into one timestamp destroys the ability to reason about it.

Attribute The question it answers Why it cannot be inferred
Source Where this value came from Two sources can disagree, and the disagreement is information
Fetch date When this value was retrieved Retrieval is not the same as the date the source itself changed
Scope Which entities and fields this value covers A source may be authoritative for names but silent on divisions

The scope attribute is the one most often missing and the most expensive to reconstruct. Without it, a blank field is ambiguous: the source said there was nothing, or the source was never asked about this field at all.

Storing the three together also makes conflicts legible. When two sources disagree about a name, a reader can see which value came from where and when, and the resolution becomes a decision with evidence rather than a coin toss.

Three update strategies and what each one costs

There are broadly three ways to keep country data current, and they trade effort against surprise in a predictable way.

Pull on a schedule is the simplest: refresh everything periodically, regardless of whether anything changed. It is easy to reason about and easy to forget, because a scheduled job that fails quietly produces no visible symptom until the data is wrong.

Pull on a signal is more efficient: watch for change notices from the bodies that publish codes and names, and refresh when something is actually announced. It reacts quickly and depends on someone reading the notices, which means it fails precisely during the periods when nobody is looking.

Refresh on demand is the least glamorous and often the most reliable for a small dataset: update a specific entry when a question about it arises, and record who asked and why. It keeps the data honest about its own gaps, at the cost of not covering entries nobody has questioned yet.

Most real systems end up combining all three, with a slow schedule as a floor, signals for the fast-moving parts, and targeted fixes for individual complaints. What matters is that the combination is deliberate and that the result is recorded per entry rather than per release.

What should a historical snapshot preserve?

A snapshot is only useful if it can answer a question about a past moment, so it has to preserve more than the values.

Preserve the values as they were, the date they described, and the source state that produced them. A snapshot that keeps only the values cannot explain why a country’s name changed between two releases, and a snapshot that keeps only the date cannot reproduce anything.

There is a limit worth setting deliberately. Keeping every field forever is expensive and rarely consulted; keeping the changed values, with attribution, is usually enough to answer the questions that actually get asked — which name was in effect for a given period, and when a code stopped being used.

The most valuable test for a snapshot is one that reconstructs a record at a past date and compares it with what the system produced at the time. If those two disagree, the snapshot is a diary rather than a record.

For developers: keep provenance on the value, not on the release

Store the source, the fetch date and the scope next to each value, or next to each group of values that shares them. Then every question about reliability becomes a query rather than a memory exercise.

When a value has no source recorded, treat that as a state worth surfacing rather than a default worth assuming. The guide to the country data coverage checklist walks through the three states a field can occupy and why an unknown should stay visible; freshness is simply the fourth dimension applied to the same table.

For the codes themselves, the published change notices are the authoritative triggers, and the guide to ISO country and subdivision codes covers how the code systems relate to one another. Keep the country and region directory as the neutral place to point readers who want to see how names and codes are presented together rather than to read about maintenance.

Source names, fetch dates and version labels used as examples above are illustrative notation only. They are not references to any particular dataset’s contents, schedule or coverage, and they describe no real update process belonging to any organisation.

Next steps

Pick one country entry and try to write down its source, its fetch date and its scope from memory. Where you cannot, the information is not stored, and the next disagreement about that entry will be settled by opinion. The guide to scaling test data across countries shows how a set of entries with uneven provenance can still be used to generate a consistent test population.

Keep reading

Address & Identity Data Formats for 86 Countries guides