Data

The self-building catalogue

Point GeoAI at a spatial database and every table becomes a searchable dataset: profiled, described in plain language, indexed by meaning and measured for where it really has data. No manual registration.

The catalogue cycle: discover a table, profile its columns, write a description, index it for search, measure coverage and overlap; human overrides and the glossary feed in; change detection re-runs the cycle
The catalogue cycle. Human corrections always win over generated text and survive regeneration. Measurements are stored per dataset and read by every later decision.

What the catalogue knows about each dataset

FactHow it is obtainedWhat it is used for
Columns and their meaningProfiled from the data: full value lists for categories, samples for varied text, minimum and maximum for numbers and dates; typed as identifier, year, location, measure, date or text.Checking every filter value; choosing sensible statistics (a year is never summed).
Ordered scalesGraded values such as energy labels or condition classes are put in order, and code checks every value appears once.“label C or better”, and the most-common value instead of a meaningless average.
Description and conceptsWritten by a model from the column evidence, in the words users use: what the dataset is, where it covers, when. Code checks it before it is saved.Semantic search, and explaining datasets to users.
ScopeYear, source, language, and whether the dataset is a subset (one year, one class, one stage of a workflow).Preferring a general inventory over a subset.
Measured coverageShapes overlaid on official municipality and neighbourhood boundaries.Coverage-aware answers.
Measured timeFor every date and year column, the span its values actually cover and how many values fall outside a plausible range, read from the data. What a column means (how current the register is, or the age of the objects) is judged by the step that reads it, not guessed by code.Telling versions of a register apart, and saying how current an answer is.
Place columnsA column counts as a place column only if every value is a real place name or code, and some lie inside the covered area. This rejects ordinary words that happen to be place names elsewhere.Filtering “in Elsrijk” on the right column, even when datasets spell names differently.
Overlap with similar datasetsMeasured both ways, on the map and by shared identifiers. Points match where they coincide; lines and areas match only when most of each lies within the other, so neighbouring areas that merely share a border are never taken for copies of each other. For features that match, what the two registers record about them is compared too: the model proposes which columns record the same attribute, from their names and sample values, and code measures how often the matched features agree. Columns that record where an object is (an address, postcode, place or district) are left out: features matched by location agree on them by definition, so they prove nothing.Deciding to combine or choose between registers.
Identifier qualityCandidate id columns are profiled; empty, constant or serial-number columns are not trusted.Linking registers only on real keys.

Search by meaning

Each description is turned into a vector for semantic search, with shared boilerplate removed so similar datasets stay distinguishable. Searches work in Dutch and English and match meaning rather than words: a question about lantaarnpalen finds the streetlight register.

Keeping it current

The catalogue compares every table with what it has stored and takes the lightest action that is enough:

Change detectedAction
New tableFull profile, description, index and measurements.
Columns added or removedDescription rewritten.
Category values or ranges changedColumn profile refreshed, no model call.
Only the row count changedCount updated.
Table emptiedMarked empty and removed from search.

This runs at start-up, once a week during off-hours, and on demand after a data load.

People stay in control

Overrides

Administrators can rewrite a dataset’s description or concepts. The correction wins everywhere, including search, and survives regeneration.

Abbreviation glossary

Codes in table names are proposed by the model and reviewed by a person. Only confirmed meanings are ever used in descriptions.

Health report

Suspected duplicates, datasets not yet measured, shapes outside the territory, and near-synonym concepts are reported for a person to decide. Nothing is hidden automatically.

Further reading on the GeoAI blog: Finding the right layer when there are two hundred.