How GeoAI works · Part 7
Finding the right layer when there are two hundred
Every GIS team has a colleague who knows what each cryptic table name means. How GeoAI learns the catalogue from the data itself, and the week every dataset looked the same.
A new colleague joins the GIS team. You spend most of their first week on one thing: explaining the database.
“The prefix is the source system. This one is the asset-management export, that one is the old viewer. The suffix is which municipality, except when it isn’t. There are three tree tables; use this one unless it’s about felling. Don’t trust the row counts on that one. The wiki page about all this is from 2019.”
By Friday they know maybe thirty of the two hundred layers. The rest they’ll learn the way you did: by asking you.
Every GIS team has this knowledge, and almost none of it is written down anywhere a machine can use. A natural-language system lives or dies on it. If someone asks about lantaarnpalen and the system can’t find the streetlight register, nothing downstream matters.
So the first thing GeoAI does, before anyone asks a question, is the same thing your new colleague does in week one. It learns the catalogue. Except it reads all two hundred layers, and it does it again every week.
Reading a table like a specialist would
For every table in the database, GeoAI builds a profile from the data itself, not from the table name:
- Columns and what they hold. Full value lists for categories, samples for free text, minimum and maximum for numbers and dates. Each column gets a type: identifier, year, location, measure, date or text.
- Ordered scales. Graded values such as energy labels or condition classes are put in order, so “label C or better” means something.
- Where it has data. Every dataset’s shapes are overlaid on the official municipality and neighbourhood boundaries, and the result is stored.
- When. For every date and year column, the span the values actually cover, so a register last updated in 2021 can be told from one updated this year.
From that evidence, a model writes a short description of what the dataset is, in the words people actually use: what it records, where, and when. Code checks the description before it’s saved.
Searching by meaning, one concept at a time
The descriptions are what GeoAI searches. They’re turned into vectors, so a search matches meaning rather than exact words, in Dutch and English. A question about lantaarnpalen lands on the streetlight register whatever its table is called.
A question usually has more than one subject: “trees near schools” is about trees and schools. Searched as one phrase, the stronger concept drowns out the weaker one and the school layer never shows up. So each concept is searched on its own, and the results are merged.
The week every dataset looked the same
Here’s a lesson that cost us time. Early descriptions ended with a helpful routing sentence, the same pattern on every dataset of a kind, telling the search what the dataset was good for. It seemed sensible.
It was a disaster for search. Datasets from the same family, say several registers of green space, ended up with descriptions that were mostly identical sentences. Their vectors became indistinguishable, and the search could no longer tell them apart. It failed silently: the right register was in the list, just tied with its siblings.
A description has one job: say what makes this dataset different from its neighbours.
The fix was to strip boilerplate from what gets embedded, and to write descriptions around identity: what this dataset is, where, and when, inferred from its columns rather than listing them. The columns themselves are already known exactly; they don’t need to be in the prose.
General inventory beats a subset
When someone asks about “trees”, a felling list for one year is technically about trees. It’s not what they mean. The catalogue records a dataset’s scope: whether it’s a subset (one year, one class, one stage of a workflow) or a general inventory. When several datasets fit, the general inventory wins, then the one with the right geography, then the most recent.
That’s the rule your new colleague learns in week two. It’s written into the ranking.
People stay in control
A catalogue written by a model will sometimes be wrong. The design assumes that:
Overrides
An administrator can rewrite any dataset’s description. The correction wins everywhere, including search, and survives regeneration.
Abbreviation glossary
Codes in table names are proposed by the model and confirmed by a person. Only confirmed meanings are ever used.
Health report
Suspected duplicates, unmeasured datasets and shapes outside the territory are listed for a person to decide. Nothing is hidden automatically.
Keeping up with a living database
Municipal databases change every week. GeoAI compares every table with what it has stored and takes the lightest action that’s enough: a new table gets the full treatment; changed columns get a new description; changed values only refresh the profile; a new row count only updates the count. That runs at start-up, weekly during off-hours, and after every data load.
The wiki from 2019 is the real competitor here. A catalogue that rebuilds itself from the data every week can’t go stale in the same way. It can be wrong, which is why a person can correct it, but it can’t silently fall behind.