LLM DEGRADATION TEST · MODEL VS MODEL · PELICAN SVG + CANDY PUZZLE
LLM Degradation Test & Model Comparison: GPT-6 vs Claude vs DeepSeek vs Kimi
Run the two probes made popular by manxue.ai — an animated pelican-on-a-bicycle SVG and a pigeonhole candy puzzle — against up to four models at once, with your own API key, and compare the answers side by side.
Real answers recorded by the maintainers with the same prompts and checks as the arena below; every SVG is re-validated in your browser before it is drawn. Scroll down, paste your own key, tick the same models and run to compare with a fresh draw.
Setup
Endpoint, key, models and probes
The key lives only in this page's memory and is sent only to the base URL above. It is never written to browser storage.
Current scene · nonce:
Each run sends one real, paid request per model per probe using your own balance. Nothing is retried. Some models reject a high output cap with HTTP 400 — lower it; that is an upstream limit, not a model failure. Claude 5 models count hidden reasoning against the cap, which is why the default is 32000 rather than manxue's 16000.
Request preview (cURL, key omitted)
Results
One card per model; badges show SVG safety, animation, nonce and the candy verdict
Pick models and run
SVG previews render as images (no scripts, no external resources); answers and sources are collapsible.
What the probes measure
The candy verdict only checks whether the answer contains the standalone number 21; it is not a semantic grade. The pelican verdict passes when the SVG parses, contains no scripts or external references, declares an animation, and shows the nonce in a <text> element. Neither probe proves model identity or overall capability. Prompts are copied verbatim from manxue-ai so results are comparable with manxue.ai.