How GeoAI works · Part 4
Why our AI isn’t allowed to write SQL
Spatial SQL doesn’t crash when it’s wrong; it returns a number. Why the model in GeoAI fills in a plan, and code writes every query.
You’ve seen the demo. Someone types a question, an AI writes SQL, the query runs, a number appears. It looks like magic.
Then you read the SQL. It joins trees to schools with a 50-metre distance condition. It runs. It returns a number. And every tree that stands near two schools has been counted twice, because nothing in the query says “each tree once”.
The number is wrong, the query is valid, and nobody in the room but you can tell.
Text-to-SQL is the obvious way to build what we built: give a language model the table definitions and let it write the query. We chose not to. Not because the models are bad at SQL. They are surprisingly good. The problem is that spatial SQL fails quietly, and it fails in exactly the places a language model is weakest.
Where spatial SQL goes wrong
Every GIS specialist carries a mental list of these:
- A join that multiplies rows, so one object is counted once per match.
- A filter on
'Basisschool'when the column actually says'basisonderwijs', which returns zero and looks like an answer. - “Inside the district” written as “touches the district”, so objects on the border of the neighbouring district sneak in.
- A year column summed instead of counted, or an average taken over energy labels.
- A district filtered by name in a table that spells it differently, or doesn’t record districts at all.
None of these raise an error. They all produce a plausible number. A language model will make each of them sometimes, and you can’t tell which time from the output.
Spatial SQL doesn’t crash when it’s wrong. It returns a number.
What the model writes instead: a plan
In GeoAI the model never writes a query. It fills in a structured plan, a form with a fixed shape. Simplified, the plan for the tree question says something like:
Simplified for this post. The real plan has more fields, but the same idea: choices, not code.
That is the part a language model is genuinely good at: understanding what you meant, and picking which dataset plays which role. It is judgement.
Everything after that is bookkeeping, and bookkeeping belongs to code.
What code does with the plan
- Checks it against the real data. Does the dataset exist? Does the column? Does
'Basisschool'actually occur in it, or is the real value spelled differently? A value that occurs nowhere is never used silently. - Picks the right spatial relation. Inside, overlapping, within a distance, further than a distance, not inside, bordering. You say which one you mean in words; code picks the matching spatial function.
- Counts each object once. Joins can’t inflate a count. An object shared by two combined registers counts once too.
- Passes values safely. Every value goes into the query as a parameter, never pasted into the SQL text.
- Checks a zero before reporting it. A zero caused by a value found nowhere in the data becomes “that value is not in the data”, not “there are none”.
The model can only choose from what the plan’s shape allows. If an analysis isn’t in GeoAI’s vocabulary of operations, the model can’t invent it, and GeoAI says it can’t do that yet. That is a limit, and we list those limits openly. It is a much better limit than a query that does something nobody asked for.
You still get to read the SQL
This isn’t about hiding the query. The opposite: every answer shows the SQL that actually ran, with “returned N of M features”. If you are the specialist in the room, you can read it the way you would read a colleague’s query, and see exactly where GeoAI’s choices differ from yours.
And the download doesn’t ask the model again. It re-runs the same verified plan without the preview limit, so the file you open in QGIS matches the answer you read, row for row.
What we gave up, and why it was worth it
A closed set of operations is less flexible than free SQL. Some questions GeoAI simply can’t answer yet, such as counts per distance band (“0–50 m, 50–100 m, …”). With free SQL, a model would have produced something for those.
That something is exactly what we don’t want. A question GeoAI can’t answer yet gets a clear “not yet”. A question it can answer gets SQL written the same way every time, by code that has been tested against a suite of real municipal questions.
The rule underneath. Judgement in the model, facts in code. What a question means and which dataset fits are judgements. Column names, values, coverage, counts and boundaries are facts, and facts are always checked by code against the real data.