Data
The self-building catalogue
Point GeoAI at a spatial database and every table becomes a searchable dataset: profiled, described in plain language, indexed by meaning and measured for where it really has data. No manual registration.
What the catalogue knows about each dataset
| Fact | How it is obtained | What it is used for |
|---|---|---|
| Columns and their meaning | Profiled from the data: full value lists for categories, samples for varied text, minimum and maximum for numbers and dates; typed as identifier, year, location, measure, date or text. | Checking every filter value; choosing sensible statistics (a year is never summed). |
| Ordered scales | Graded values such as energy labels or condition classes are put in order, and code checks every value appears once. | “label C or better”, and the most-common value instead of a meaningless average. |
| Description and concepts | Written by a model from the column evidence, in the words users use: what the dataset is, where it covers, when. Code checks it before it is saved. | Semantic search, and explaining datasets to users. |
| Scope | Year, source, language, and whether the dataset is a subset (one year, one class, one stage of a workflow). | Preferring a general inventory over a subset. |
| Measured coverage | Shapes overlaid on official municipality and neighbourhood boundaries. | Coverage-aware answers. |
| Measured time | For every date and year column, the span its values actually cover and how many values fall outside a plausible range, read from the data. What a column means (how current the register is, or the age of the objects) is judged by the step that reads it, not guessed by code. | Telling versions of a register apart, and saying how current an answer is. |
| Place columns | A column counts as a place column only if every value is a real place name or code, and some lie inside the covered area. This rejects ordinary words that happen to be place names elsewhere. | Filtering “in Elsrijk” on the right column, even when datasets spell names differently. |
| Overlap with similar datasets | Measured both ways, on the map and by shared identifiers. Points match where they coincide; lines and areas match only when most of each lies within the other, so neighbouring areas that merely share a border are never taken for copies of each other. For features that match, what the two registers record about them is compared too: the model proposes which columns record the same attribute, from their names and sample values, and code measures how often the matched features agree. Columns that record where an object is (an address, postcode, place or district) are left out: features matched by location agree on them by definition, so they prove nothing. | Deciding to combine or choose between registers. |
| Identifier quality | Candidate id columns are profiled; empty, constant or serial-number columns are not trusted. | Linking registers only on real keys. |
Search by meaning
Each description is turned into a vector for semantic search, with shared boilerplate removed so similar datasets stay distinguishable. Searches work in Dutch and English and match meaning rather than words: a question about lantaarnpalen finds the streetlight register.
Keeping it current
The catalogue compares every table with what it has stored and takes the lightest action that is enough:
| Change detected | Action |
|---|---|
| New table | Full profile, description, index and measurements. |
| Columns added or removed | Description rewritten. |
| Category values or ranges changed | Column profile refreshed, no model call. |
| Only the row count changed | Count updated. |
| Table emptied | Marked empty and removed from search. |
This runs at start-up, once a week during off-hours, and on demand after a data load.
People stay in control
Overrides
Administrators can rewrite a dataset’s description or concepts. The correction wins everywhere, including search, and survives regeneration.
Abbreviation glossary
Codes in table names are proposed by the model and reviewed by a person. Only confirmed meanings are ever used in descriptions.
Health report
Suspected duplicates, datasets not yet measured, shapes outside the territory, and near-synonym concepts are reported for a person to decide. Nothing is hidden automatically.
Further reading on the GeoAI blog: Finding the right layer when there are two hundred.