12 KiB
How this skill stays close to the SAP standard
A claim like "this plugin produces SAP-Architecture-Center-style diagrams" is only believable if there's an empirical way to measure it. This file documents the comparison harness, the fidelity numbers, and the workflow that produces high-fidelity output.
The fingerprinting harness — scripts/compare.py
compare.py extracts a structural + style fingerprint from any .drawio file and computes a similarity score against another .drawio file. The fingerprint covers:
| Dimension | What's checked |
|---|---|
| Canvas | pageWidth × pageHeight — should match the selected SAP template; 1169 × 827 is the default for new L2 diagrams |
| Page background | pageBackgroundColor / background attribute — SAP diagrams use white/transparent. A non-white candidate scores 0 on this metric. |
| Counts | total cells, vertices, edges, inline-SVG icons, legacy mxgraph.sap.icon stencil count, pills (arcSize=50) |
| Zone hierarchy | nesting depth of zone cells — catches structural mistakes like nesting a focus zone inside another one when the SAP reference puts them side by side (Joule-inside-BTP bug) |
| Palette | the set of hex colors in the file (Jaccard similarity) |
| Edge palette | the set of strokeColor values actually used on edges — catches semantic color swaps (green↔magenta) that the global palette set hides |
| Pill vocabulary | how many pills use canonical SAP verbs (TRUST/Authenticate/A2A/MCP/ORD/HTTPS/OData/REST/SAML2/OIDC/...) vs novelty verbs (PROMPT/ROUTE/CONTEXT/...) |
| Fonts | fontFamily values used (subset = full credit) |
| Stroke widths | the set of strokeWidth values |
| Polish | presence of absoluteArcSize=1, labelBackgroundColor=default, grid-snap rate |
| Labels | visible label count and label-token overlap, so wrong-target templates no longer score as perfect |
The score is a weighted blend of these dimensions; 100 means the two files have an identical fingerprint, 0 means nothing in common.
python3 scripts/compare.py reference.drawio candidate.drawio
python3 scripts/compare.py --score reference.drawio candidate.drawio # one-line score
python3 scripts/compare.py --json reference.drawio candidate.drawio # machine-readable
Calibration:
| Pair | Expected score |
|---|---|
| File compared to itself | 100 |
| Different SAP-published L2 diagrams (e.g. IAS Authentication vs Task Center) | 80–85 |
| L0 of a scenario vs L2 of the same scenario | 60–70 |
| Hand-crafted candidate built from scratch | 50–55 |
| Candidate built by copying a reference + relabeling (the recommended workflow) | 95–100 when the target scenario stays close |
The big gap between "from-scratch" (≈50) and "from-template" (≈100) is the empirical justification for the SKILL.md rule: never draw from scratch — always start from a reference template.
The full quality loop
description ┐
│
▼
┌──────────────────────────────────────────────────────────┐
│ Step 1 — scaffold from a SAP reference template │
│ scaffold_diagram.py "<request>" --out <file>.drawio │
│ Ranks 71 bundled SAP templates and copies the best one │
│ (uses metadata aliases/tags + visible draw.io labels │
│ + the "primary": true flag for canonical umbrella refs)│
└──────────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────┐
│ Step 2 — surgical relabel for the new scenario │
│ Title, zone labels, service-card values; preserve │
│ canvas size, zone hierarchy, edges, pills, legend, │
│ network divider, SAP logos, footer │
└──────────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────┐
│ Step 3 — autofix.py --write │
│ Snap grid, normalise hex case, fix arcSize, strokeWidth│
└──────────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────┐
│ Step 4 — validate.py │
│ Errors: bent arrows, label overflow, sibling overlap, │
│ missing geometry, duplicate ids, │
│ dark/branded page background │
│ Warnings: off-palette, off-grid, missing label-bg, │
│ off-vocabulary pill verbs, multi-logo over-use │
└──────────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────┐
│ Step 5 — score_corpus.py across all bundled references │
│ Best score should be ≥ 90 if template drift is low │
│ If < 90, compare.py shows where the structure drifted │
│ (canvas, page bg, zone depth, edge palette, pill vocab)│
└──────────────────────────────────────────────────────────┘
│
▼
final .drawio + flow narration
For eval_corpus.py run --exclude-target-template, the exact target is removed from the selector pool. The harness therefore adds an explicit primary visual-neighbor hint computed with compare.py fingerprints. This is not used for normal production generation; it makes the leave-one-out research loop test the closest available SAP layout instead of an arbitrary semantic neighbor.
The harness also records the selected-template target baseline for every case. This is the score of the alternate template before any model label plan is applied. In overnight leave-one-out runs, low baseline scores are a ceiling signal: label edits can improve content overlap, but they cannot invent the target's canvas rhythm, vertex count, zone proportions, edge topology, or service-icon density. Reports classify these as:
near-miss: failed, but within--retry-marginof--min-score; retrying, adding metadata, or improving label replacement may help.ceiling-limited: failed below the retry floor; add a closer SAP sibling template or implement geometry-aware generation before spending more model time.model-failure/validator-failure: fix the generation or validation error first.
The default --retry-margin 8 means a --min-score 90 run stops retrying cases below 82. This reflects the observed overnight loop: most large gaps were template-coverage gaps, not stochastic model failures.
Worked example — examples/iam-arc1-mcp-l2.drawio
The bundled examples/iam-arc1-mcp-l2.drawio was produced by:
- Picking
btp_SAP_Cloud_Identity_Services_Authentication_L2.drawioas the closest reference template (it's the canonical IAM-on-BTP diagram). - Surgically swapping ~5 labels for an ARC-1 MCP scenario:
- title:
Authentication with SAP Cloud Identity Services→ARC-1 MCP Server - Authentication on SAP BTP - subtitle:
Recommended authentication flows…→Claude Desktop / Copilot Studio MCP clients calling ARC-1 over XSUAA OAuth, reaching on-prem SAP via Cloud Connector with Principal Propagation - card label:
SuccessFactors→ARC-1 MCP Server - zone label:
SAP BTP Applications -IAM based on SAP Cloud Identity Services→SAP BTP Applications - ARC-1 MCP based on SAP Cloud Identity Services - card label:
Mobile/Desktop→Claude Desktop / Copilot Studio
- title:
- Running
autofix.py --write(resulted in 436 mechanical fixes — geometry snap, hex case, arc size, font normalisation, comment strip). - Running
validate.py— exit 0. - Running
compare.pyagainst the original reference — scored 96.6/100 with the target-aware label-token scorer. - Running
score_corpus.py --min-score 90across the bundled templates — best score 96.6/100.
This proves the workflow: with a few hand-edits, you preserve SAP's visual structure while the scorer still notices intentional scenario-label changes.
Why the validator + autofix matter
Without these gates, a hand-crafted candidate scored ~52 even when it followed the rules in references/. The biggest contributors to the gap:
- Bent
orthogonalEdgeStylearrows (centers not aligned) - Sparse zones with too few service cards (low vertex / icon count)
- Off-palette hex from improvising "close-enough" colors
- Missing
labelBackgroundColor=defaulton edge labels - Dark / branded page background (now a hard validator error)
- Off-vocabulary pill verbs like
PROMPT,ROUTE,CONTEXT,DELEGATE— replaced withTRUST,Authenticate,A2A,MCP,ORD,HTTPS,OData/REST,SAML2/OIDC - Wrong zone hierarchy (e.g. nesting Joule inside the BTP zone when SAP places it as a sibling) — now penalised by the
zone_depthmetric incompare.py
validate.py catches all of these before the diagram is shown to the user. autofix.py repairs the mechanical ones automatically.
Limitations
The fingerprint compares structure, style, and visible label overlap — not full semantic correctness. Two diagrams with similar fingerprints can still encode different architectures. The score validates "looks SAP-styled and uses similar target labels" but doesn't validate "the architecture actually works".
Also, the validator can't check:
- One-SAP-logo rule — multiple logos count as warnings only when they appear inline in the XML
- Semantic correctness of arrows — green / pink / indigo edges are colored correctly, but whether that specific edge should be authentication, trust, or authorization is a judgment call left to the author
- Legend completeness — presence of a legend block is checked; whether it accurately covers all colors in the diagram is not
For those, manual review against references/do-and-dont.md remains necessary.
Corpus scoring
score_corpus.py wraps compare.py and ranks the candidate against every bundled .drawio reference:
python3 scripts/score_corpus.py --top 5 --min-score 90 my-diagram.drawio
Use this as the final fidelity gate. A good template-derived diagram should have:
| Signal | Target |
|---|---|
| Best target/corpus score | >= 90 |
| Chosen-template pairwise score | >= 90, ideally 95-100 |
| Validator errors | 0 |
| Off-palette / line-style drift | explainable or fixed |
For research runs against SAP's full public corpus, clone the upstream repositories and pass them as reference directories:
python3 scripts/score_corpus.py \
--references /path/to/SAP/btp-solution-diagrams \
--references /path/to/SAP/architecture-center \
my-diagram.drawio
See corpus-findings.md for the 2026 snapshot that motivated the current 71-template bundle.
How to add new reference templates
- Drop a
.drawiofile inassets/reference-examples/(any name) - Confirm it's Apache-2.0 / MIT / your own work
- Add an entry to
assets/NOTICE.mdif the source is third-party - Re-score your test diagrams against the new template —
score_corpus.py --top 10
The skill picks the highest-scoring reference automatically when the user describes a scenario, so adding more references improves quality monotonically.