Projects05 of 15
Vi-Graph
Diagram images turned into editable, queryable graphs
The question
How accurately can a compact vision-language model recover the structure of a diagram, and how does that change as diagrams get more complex?
What it does
Vi-Graph takes a picture of an architecture, an ML pipeline, a flowchart or a system diagram and rebuilds it as a strict, versioned JSON graph. The graph can then be:
- edited in the browser;
- exported to Mermaid, SVG, PNG, PDF or JSON;
- asked questions, such as what feeds into this or which branches run in parallel. The answer comes from the graph and highlights the nodes it rests on.
The model
Qwen3-VL-2B-Instruct, fine-tuned with 4-bit QLoRA: 17.4M trainable parameters, 250 steps, 1.4 hours and 7.6 GB of peak GPU memory.
A benchmark built for the question
Real diagrams rarely come with exact structural ground truth, so the benchmark was built:
- 2,800 rendered diagrams across seven types and four difficulty levels, from 3–6 nodes up to 25–40;
- a test split whose layouts and themes never appear in training;
- 25,698 structural questions generated from the ground truth.
Results
On all 500 held-out test diagrams:
| Measure | Before fine-tuning | After |
|---|---|---|
| Graph similarity | 0.581 | 0.740 |
| Edge F1 | 0.499 | 0.638 |
| Structural QA | 0.437 | 0.582 |
| Valid graphs | 72% | 91% |
Every metric improves at p < 0.001 under a paired bootstrap with Holm correction, and the result is stable across sampling seeds at 0.748 ± 0.006.
The baseline that nearly wins
The baseline that might have beaten it was built too, and it nearly does. A classical OCR and OpenCV pipeline ties it on graph similarity, 0.738 against 0.740, and beats it on node F1 and on validity.
The model leads on edges, question answering and diagram type. It wins on small diagrams and falls behind on the densest.
Why the densest diagrams are harder
On 25–40-node diagrams the model sometimes keeps listing invented nodes until it runs out of tokens: runaway enumeration. Where it does return a graph, it matches the baseline.
- An inference-time guard that stops those loops cut evaluation time by 26%, with no significant change in any score.
- Retry-and-repair turns 72.8% usable first answers into 91.0%.
- Resolution has a floor at 640 px; 768, 896 and 1024 score the same.
Traceable by design
Every inference logs the model revision, adapter, prompt hash, decoding parameters, seed, schema version and split hash. There are 767 automated tests.
Limits
All results are on synthetic diagrams, with one model family and size, and the label-matching threshold is not yet calibrated on hand-labelled real data.
How it was built
PyTorch, Transformers, PEFT/QLoRA, FastAPI, Next.js and React Flow. Built with Claude Code as a pair programmer; designed and directed by me, with all training and GPU evaluation run on my own hardware. The demo’s model outputs are replayed from the recorded evaluation run; everything after the model runs live.