LLM DEGRADATION TEST · MODEL VS MODEL · PELICAN SVG + CANDY PUZZLE

LLM Degradation Test & Model Comparison: GPT-6 vs Claude vs DeepSeek vs Kimi

Run the two probes made popular by manxue.ai — an animated pelican-on-a-bicycle SVG and a pigeonhole candy puzzle — against up to four models at once, with your own API key, and compare the answers side by side.

Setup

Endpoint, key, models and probes

REQUEST

The key lives only in this page's memory and is sent only to the base URL above. It is never written to browser storage.

Current scene · nonce:

Each run sends one real, paid request per model per probe using your own balance. Nothing is retried. Some models reject a high output cap with HTTP 400 — lower it; that is an upstream limit, not a model failure. Claude 5 models count hidden reasoning against the cap, which is why the default is 32000 rather than manxue's 16000.

Request preview (cURL, key omitted)

Results

One card per model; badges show SVG safety, animation, nonce and the candy verdict

RESULTS

Pick models and run

SVG previews render as images (no scripts, no external resources); answers and sources are collapsible.

What the probes measure

The candy verdict only checks whether the answer contains the standalone number 21; it is not a semantic grade. The pelican verdict passes when the SVG parses, contains no scripts or external references, declares an animation, and shows the nonce in a <text> element. Neither probe proves model identity or overall capability. Prompts are copied verbatim from manxue-ai so results are comparable with manxue.ai.