How GeoAI works · Part 10
A layer name is a promise the data doesn’t have to keep
A suffix everyone read as a region turned out to mean one town. Why GeoAI measures coverage instead of reading it from names, and why an empty field beats a confident guess.
For a while, part of our system believed a particular three-letter suffix on table names meant a whole region: several municipalities working together. It was a reasonable guess. The letters matched the abbreviation of a well-known regional partnership. A model had proposed it, it sounded right, and it quietly spread into descriptions, search vectors and our own notes.
Then someone ran a geometry-extent check on the layers carrying it. Every feature sat inside one single municipality.
The suffix didn’t mean a region. It meant one town.
Every GIS specialist has learned this one, usually the hard way: a layer’s name is a claim, not a fact. The name records what someone intended when they created the table. The data records what actually happened afterwards. They drift apart constantly.
Three ways names lie
The suffix that means something else
A code everyone “knows” turns out to mean something narrower, or broader, than assumed. Nobody wrote it down, so the guess becomes the truth.
The municipal layer that isn’t
A layer named for one municipality holds data for two. Another, presented as general, holds almost everything for one municipality and a thin scattering for the other.
The national register that spills
An extract of a national register includes a band of neighbouring municipalities. “All buildings” quietly includes buildings that aren’t yours.
In our sample data, one green-space layer that looked general held around four fifths of its records in one municipality. Anyone answering “how much green space is there per district?” from that layer, without knowing, would conclude the other municipality is barely green.
The name tells you what someone intended. Only the geometry tells you what’s there.
Measure, don’t read
So GeoAI doesn’t read coverage from names at all. For every dataset, the catalogue overlays its shapes on the official municipality and neighbourhood boundaries and stores how many features fall in each. It separates meaningful coverage from spill: a few features lying just across a border don’t make the dataset a register for that neighbour.
Coverage is re-measured on every catalogue sync. A dataset that has never been measured is treated as unknown, never as “covers nothing”.
Judge only the data you have
Measured coverage matters most for a kind of question GIS specialists handle with great care, and naive systems handle badly: absence.
“Which homes are more than 500 m from a waste container?” sounds simple. But if the container register only records one municipality, every home in the other one is “far from a container”, because nothing was ever recorded there. The answer would be large, confident and meaningless.
GeoAI restricts every absence question to where the other register actually has data, and the scope note says so. The same principle runs through everything:
- “Areas with none of X” only considers areas where X is recorded at all.
- A heavily lopsided register gets a warning when you ask about the place it barely covers: “records N here against M there; not a complete inventory”.
- With no place in the question, national data is restricted to your territory, and the answer names what was left out.
- A ratio keeps its denominator to the area the counted dataset actually covers.
We summarise it as one sentence, and it has become a rule for the whole project: judge only the data we have, and be exact about it. A missing register is a data gap for a person to fix. It is never a licence for the system to fill in the blank.
Abstain beats guess
The suffix story taught us a second rule, maybe the more important one. The mistake wasn’t that a model guessed. Models guess. The mistake was that the guess was stored, and everything downstream treated it as fact.
Now, when the catalogue can’t ground a fact (a year, a source, what an abbreviation means, what area a layer covers), it leaves the field empty. A missing fact is recoverable: someone notices, someone fills it in. A confident wrong fact steers every answer that touches it, silently, for weeks.
That’s also why abbreviations in table names now go through a glossary: the model proposes a meaning with its evidence, and a person confirms it before it’s ever used.
We still find holes. Recently, an absence question over a national extract treated the extract’s spill into a neighbouring city as coverage, and a hundred-odd districts there came back as having “no bins”. The rule was right; one path didn’t apply it. It’s exactly the kind of mistake this post is about, and we found it the same way: by looking at where the answer’s data actually lay, instead of trusting what the plan said.