Keystone vs Hallmark · comparison gallery
Same brief verbatim into both skills. Same model, no human intervention.
Both outputs rendered by Keystone's Playwright at 5 viewports and scored by
Keystone's engine — the same engine that scores Keystone's own output.
Losses are published (the honesty clause). Vision S1 rows
(“does this look AI-generated?”) are appended per brief in verdict.md.
Generated by test/compare/run-comparison.mjs · Plan 5b.
| Side | Engine score /48 | Failed gates | Open |
|---|