Start here
How it works
Every message passes through the same eight stages. Language models are used where judgement is needed, such as reading a question or choosing a dataset. Every fact, count and query comes from deterministic code that checks the model’s work.
Stage by stage
| Stage | What happens | What you see |
|---|---|---|
| 1 Route | The message is sorted into chat, a question about the catalogue, or a spatial question. Its language is noted. A question asking for two different kinds of thing is split into parts, each written as a complete question with its own places and conditions. | Nothing for chat; otherwise the first steps in the thinking panel. |
| 2 Understand | The question is broken into concepts, named places, grouping (per district, neighbourhood or municipality), time conditions and required conditions. Places are matched to official codes by fixed rules. | “Recognised places: …” |
| 3 Find data | When a place is named, only datasets with data in that municipality are searched. Each concept is then searched separately by meaning, in Dutch and English, and the results are merged. | Candidate datasets with scores |
| 4 Select | Datasets get roles: the target that is counted, constraints that restrict it, and a join partner it is linked to. District boundaries are attached when needed. | “Selected: …” with a reason |
| 5 Dataset | Related registers of the same kind in the asked area, including ones the search did not rank, are combined or one is chosen, based on measured overlap. Each condition is bound to a real column and a real value. | Shown on the Stage |
| 6 Plan | A structured plan is written, never SQL. A per-area question that comes back grouped on a missing column is asked again once, to group by the official boundaries. More than thirty checks then verify every name, value, threshold and place, repair what can be repaired (for example, holding a per-area plan to the named place), and stop what cannot. | The plan, or a clarifying question |
| 7 Execute | Code turns the plan into a safe database query and runs it. A suspicious zero is checked before it is reported. | The query and “returned N of M” |
| 8 Explain | The answer is written from the actual rows and the executed query, with a scope note stating what was counted and where. | The streamed answer and downloads |
Two design rules that shape everything
Judgement in the model, facts in code
What a question means and which dataset fits are judgements. Column names, values, coverage, counts and place boundaries are facts, and facts are always checked by code against the real data.
Repeatable decisions
Decision steps use fixed settings with no random sampling, which keeps plans reproducible and lets every change be tested against a suite of real municipal questions before it ships.
Further reading on the GeoAI blog: One question, eight stops and Why our AI isn’t allowed to write SQL.