From our research programme

Selected research fields for AI agents in technical software.

Our research programme studies how AI agents perform domain tasks, how their results can be evaluated methodically and how AI can be integrated deeply into GIS and CAD systems. This page presents selected priorities: SpatialAgents benchmarks, CAD-Agents, SpatialApp and Holistech Labs.

SpatialAgents benchmarks · Benchmark 08

One QGIS workflow, seven agent-model systems.

Benchmark 08 evaluates a multi-step climate and hazard analysis for a study area in Stuttgart.

The SpatialAgents in each run execute the workflow inside QGIS and produce structured project, data and map artefacts. Seven agent-model systems work on the same task in three runs each, six of them in the cloud and one locally on a notebook. Further SpatialAgents benchmarks will extend this field of research.

Open the complete benchmark record
QGIS project from SpatialAgents Benchmark 08 with risk layers, buildings and evacuation routes
Benchmark 08 / QGIS projectThe QGIS project from a valid benchmark run with the generated risk and evacuation layers.

Results from Benchmark 08.

Benchmark 08 results, points out of 100
ModelAgentExecutionRun 1Run 2Run 3MeanBand
Claude Opus 5Claude CodeCloud92919091.0Gold
Claude Sonnet 5Claude CodeCloud85858685.3Gold
Qwen3.8-Flash-NextOpenCodelocal · HP ZBook81838783.7Silver
GPT-5.6 SolCodex CLICloud78838381.3Silver
GPT-5.6 TerraCodex CLICloud78757977.3Silver
GPT-5.6 LunaCodex CLICloud71727572.7Silver
Claude Haiku 4.5Claude CodeCloud53595254.7Passed

The means are calculated from the three runs shown for each model. Twenty-one runs were included; four scored runs were excluded and three runs failed for environmental reasons.

Qwen3.8-Flash-Next ran locally on an HP ZBook Ultra G1a and reaches 83.7 points — more than any of the three GPT-5.6 variants in the cloud. In local operation no request leaves the organisation's own network. The same machine carries our inference-engine measurements in the Holistech Labs section.

SpatialAgents benchmarks · Alpen case file

Eight public-authority tasks on open ALKIS data.

The second measurement series covers B20 to B27, tasks from a municipal administration: exposure within the flood area, distances to housing and protected areas, area balance, green share, reprojection and geometry repair.

34runs
scoredacross 8 tasks and two models
17 of 17gold
Qwen3.8-Flash-Nextlocally on a notebook
16 of 17gold
Claude Sonnet 5in the cloud, one run failed
10runs
invalidkept separate, not scored

Every check names its evidence and every map image is documented. The complete case file is published on spatialagents.de in German.

Open the case file

CAD-Agents

Five dimensions for CAD reconstruction.

The CAD testbed measures dimensional accuracy, feature fidelity, shape fidelity, drawing fidelity and thread fidelity where applicable.

Open the research site
Overall CAD-Agents workflow: technical drawing, AI agent, CAD build with FreeCAD, delivery, evaluation and results matrix
CAD-Agents / Overall workflowFrom the technical drawing through the AI agent and FreeCAD to evaluation and the results matrix.

SpatialApp · deep AI integration

Research GIS applications and SpatialAgents as one system.

With SpatialApp, we research how SpatialAgents can be integrated deeply into GIS applications.

SpatialAgents can capture geodata and application state in structured form, execute GIS functions, operate user interfaces and verify the effect of their work. Visual guidance, reproducible workflows and specialised GIS applications form part of the same platform.

Explore Düngewende as a specialist application ↗
SpatialApp example project with thematic risk layers and the Hochrisiko bookmark in Basel
SpatialApp / Example projectThematic risk layers and the “Hochrisiko” bookmark in Basel.
SpatialAgents guide inside SpatialApp at the Hochrisiko area in Basel
SpatialApp / SpatialAgentsThe Interaction API connects SpatialAgents with visible guidance and the defined geographic extent.

Holistech Labs

Measurements you can recheck.

Our measurement lab tests hardware, inference engines and models under controlled conditions and publishes the run logs. Two series so far, both on the same mobile workstation. The same machine solves the domain task in Benchmark 08.

Measured 8 September 2026

llama.cpp against halogen-flash-server.

The same model, the same twenty tasks, the same order: 17 against 18 tasks solved, but 486 against 157 minutes.

Dot plot of the runtime of all twenty tasks on a logarithmic scale from 30 seconds to more than 100 minutes: one dot per task for llama.cpp and one for halogen, with halogen consistently further left and therefore faster; open circles mark tasks that were not solved
Report 01 / Runtime per taskEach of the twenty tasks on its own: the gap runs through the whole set rather than resting on a few outliers.

Qwen3.8-Flash-Next on the HP ZBook Ultra G1a

886tokens/s
Prompt processingon a 69,000-token context
39.9tokens/s
Output on long contextthe same 69,000 tokens ahead of it
39.3tokens/s
Output in chatshort questions, under 800 prompt tokens
157minutes
for all twenty tasks339,134 tokens generated

Measured with halogen-flash-server. Means over the tasks in each group, measured on the run's own answers, not on filler prompts. Temperature 0, prompt cache off, 70 W GPU budget.

Open the measurement report

Measured 15 June 2025

Nine models on the same machine.

From the 0.6-billion model to more than 100 billion: output speed, wait and reproducibility across three runs each. A factor of 9.5 separates the fastest model from the slowest.

Open the measurement report