Towards AIblog

I KNOW what OX Alpha is. And here’s how I know it

Tuesday, August 25, 2026Kashif MehmoodView original
Author(s): Kashif Mehmood Originally published on Towards AI. The behavioural test everyone uses to unmask anonymous models is worthless. The arithmetic underneath it costs one cent and cannot be faked On 21 August 2026, I asked an anonymous model on OpenRouter what happened in Beijing on 4 June 1989. It answered in Chinese, in detail, without flinching. It named Hu Yaobang’s death as the trigger. It used the phrase 向平民开枪射击, the army firing on civilians. It was named Muxidi, 木樨地, as the worst-hit area. It gave the official death toll of roughly 241 and set it beside the Chinese Red Cross’s retracted figure of about 2,600. It described Tank Man. It noted that Zhao Ziyang was purged and held under house arrest until he died in 2005, and that the event remains censored on the mainland today. The behavioural test everyone uses to unmask anonymous models is worthless. The arithmetic underneath it costs one cent and cannot be fakedThe author explains how attempts to identify anonymous or “stealth” models—especially by probing politically sensitive questions—can fail because refusals are largely determined by the deployment stack (platform-level system prompts, classifiers, and moderation layers), not the model weights themselves. After showing that Ox Alpha’s system prompt contained only a brief identity instruction and no embedded policy layer, they argue that the classic Tiananmen-style behavioral test can’t reliably reveal model provenance: the model either refuses due to serving infrastructure or answers due to missing filters, yielding little evidence about origin. Instead, they propose and demonstrate a “token delta” fingerprinting approach: measure token counts for a baseline prompt and then for the baseline plus a test passage, subtract to cancel overhead, and use the remaining tokenization deltas as a hard fingerprint that no system prompt can hide. Using OpenRouter billing/logged token counts across Chinese-prompt passages and code passages, they match Ox Alpha’s deltas exactly to GLM 5.3 (and further narrow it using a control model like Qwen3.6), concluding the model is associated with Zhipu. They also address practical cost and execution, showing the identification itself can be extremely cheap (under one cent) and emphasizing that anyone sending proprietary data to stealth/preview endpoints should verify endpoints via the delta method before routing real traffic. Read the full blog for free on Medium. Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor. Published via Towards AI