giobuilds ← All posts ← Todos los posts

AI models

Ox Alpha: what the full run actually said

Ox Alpha: lo que el run completo sí dijo

On August 21 Ben Davis ran Ox Alpha on 10 DeepSWE tasks and solved 8. An anonymous model, free, with an 80 percent resolve rate. The headline wrote itself.

On the 22nd, Henry Zhang ran all 113 tasks. 66 solved. 58.4 percent.

Still a huge result, but not 80. The 80 was a subset. The 58.4 doesn't make for pretty headlines.

What Ox Alpha is

A reasoning model released in stealth mode on August 20 through OpenRouter, identifier stealth/ox-alpha. Nobody knows who made it. OpenRouter only routes requests and clarifies it is not the developer or the owner.

The rest is documented:

  • Context window of 1,048,576 tokens, max output of 131,072.
  • Accepts text, images and video. Returns text only.
  • Reasoning always on, tool calling, structured JSON output.
  • Price: zero. Free preview.

Patrick Collison, CEO of Stripe, publicly called it "very impressive". That kind of thing fuels an anonymous launch better than any marketing department.

The forensics, the fun part

Nobody has claimed it, but the evidence points one way.

Joseph W. Elstner (implicator.ai) ran 95 tokenizer probes against public vocabularies. All 95 match GLM-5 from Zhipu exactly. Mean absolute error: 0.00. Not "similar". The same vocabulary.

There's more. Error stack traces show Java classes like com.wd.paas.api.domain.v4.chat.ChatCompletionRequest and error code 1214, which match Z.ai's documented infrastructure. Video token budgets match GLM-5V-Turbo. The style of generated code, the audio refusals and the chat templates also match the GLM family.

And there's precedent. Previous OpenRouter "Alphas" ended up claimed by Chinese labs: Pony Alpha was GLM-5, Hunter Alpha was Xiaomi's MiMo.

Dominant hypothesis: Z.ai / Zhipu, some variant of GLM-5.3 or a multimodal cousin. Official confirmation: none. Andrew Curran and other analysts had doubts on the 22nd and 23rd, and someone mentioned Microsoft MAI, but the hardest technical evidence keeps winning.

The numbers, seriously now

Full DeepSWE: 58.4 percent, practically tied with Claude Opus 4.8, which sits around 59. Details that matter more than the round number:

  • In 80 percent of tasks it solved at least 90 percent of the fail-to-pass tests.
  • 9.7 percent of the failures were tool-call format, not reasoning: the model knew how to solve the task and lost points on how it called tools.
  • Secondary reports put it around 63 percent.

On Kingbench it scores 87.5, behind the published GLM-5.3 at 91.25. Real-world feedback is mixed: some find it excellent and very smooth, others see it lose in head-to-head comparisons against Kimi K3. The early hype oversized it on computer-use and self-steering.

The adoption is real

This isn't measured in likes. On OpenCode, in three days: 16 trillion tokens processed, 221,000 unique users, more than 5 million sessions. Number 2 by recent usage there. On OpenRouter the volume went from 1.5 trillion to more than 6 trillion tokens between August 20 and 23. The top app using it, Hermes Agent, pushed more than 2 trillion on its own. Almost all of the consumption is coding agents.

99.99 percent uptime over three days, 88 percent cache hit, P50 latency of 5.2 seconds. Whoever runs this knows how to serve.

My rule while it lasts

The free window is estimated to run until August 27: OpenCode promised "free for the next week" from launch, and OpenRouter publishes no cutoff date.

OpenRouter says the provider retains prompts and completions, though it claims not to use them for training. OpenCode declares zero retention on its route. I wouldn't send corporate code, or anything you don't want to see in someone else's log, while the origin stays unconfirmed. Free isn't the same as cheap: here the price is your prompts feeding the biggest public eval of the moment.

Use it for prototypes and to learn where the frontier is. Work code stays home.

El 21 de agosto Ben Davis corrió Ox Alpha en 10 tareas de DeepSWE y resolvió 8. Un modelo anónimo, gratis, con 80 por ciento de resolved rate. El titular se escribió solo.

El 22, Henry Zhang corrió las 113 tareas completas. 66 resueltas. 58,4 por ciento.

Sigue siendo un resultado enorme, pero no es el 80. El 80 era un subset. El 58,4 no da titulares tan bonitos.

Qué es Ox Alpha

Un modelo de razonamiento publicado en modo stealth el 20 de agosto en OpenRouter, con identificador stealth/ox-alpha. Nadie sabe quién lo hizo. OpenRouter solo enruta peticiones y aclara que no es el desarrollador ni el dueño.

Lo demás sí está documentado:

  • Ventana de contexto de 1.048.576 tokens, salida máxima de 131.072.
  • Acepta texto, imágenes y video. Devuelve solo texto.
  • Razonamiento siempre activo, tool calling, salida estructurada en JSON.
  • Precio: cero. Preview gratuito.

Patrick Collison, CEO de Stripe, lo llamó "very impressive" en público. Ese tipo de cosas alimenta un lanzamiento anónimo mejor que cualquier departamento de marketing.

La parte forense, que es la divertida

Nadie lo ha reclamado, pero la evidencia apunta en una sola dirección.

Joseph W. Elstner (implicator.ai) corrió 95 probes de tokenizer contra vocabularios públicos. Los 95 coinciden exactamente con GLM-5 de Zhipu. Error absoluto medio: 0.00. No es "parecido". Es el mismo vocabulario.

Hay más. Los stack traces de errores muestran clases Java como com.wd.paas.api.domain.v4.chat.ChatCompletionRequest y un código de error 1214, que calzan con la infraestructura documentada de Z.ai. Los presupuestos de tokens de video coinciden con GLM-5V-Turbo. El estilo del código generado, los rechazos de audio y las plantillas de chat también calzan con la familia GLM.

Y hay precedente. Los "Alphas" anteriores de OpenRouter terminaron reclamados por labs chinos: Pony Alpha era GLM-5, Hunter Alpha era el MiMo de Xiaomi.

Hipótesis dominante: Z.ai / Zhipu, alguna variante de GLM-5.3 o un primo multimodal. Confirmación oficial: ninguna. Andrew Curran y otros analistas tuvieron dudas el 22 y 23, y alguien mencionó a Microsoft MAI, pero las pruebas técnicas más duras siguen ganando.

Los números, ya en serio

DeepSWE completo: 58,4 por ciento, prácticamente empatado con Claude Opus 4.8, que ronda el 59. Detalles que importan más que el número redondo:

  • El 80 por ciento de las tareas resolvió al menos 90 por ciento de los tests fail-to-pass.
  • El 9,7 por ciento de los fallos fue por formato de tool-call, no por razonamiento: el modelo sabía resolver la tarea y perdió puntos por cómo llamaba a las herramientas.
  • Reportes secundarios lo citan alrededor de 63 por ciento.

En Kingbench da 87,5, detrás del GLM-5.3 publicado (91,25). En uso real el feedback es mixto: hay quien lo encuentra excelente y muy suave, y quien en comparaciones directas contra Kimi K3 lo ve perder. El hype de los primeros días le quedó grande en computer-use y self-steering.

La adopción sí es real

Esto no se mide en likes. En OpenCode, en tres días: 16 billones de tokens procesados, 221.000 usuarios únicos, más de 5 millones de sesiones. Es el número 2 por uso reciente ahí. En OpenRouter el volumen pasó de 1,5 billones a más de 6 billones de tokens entre el 20 y el 23 de agosto. La app que más lo usa, Hermes Agent, le metió más de 2 billones sola. Casi todo el consumo son agentes de código.

Uptime de 99,99 por ciento en tres días, cache hit del 88 por ciento, latencia P50 de 5,2 segundos. Quien sea que lo opera, sabe servir.

Mi regla mientras dure

La ventana gratis se estima hasta el 27 de agosto: OpenCode prometió "free for the next week" desde el lanzamiento, y OpenRouter no publica fecha de corte.

OpenRouter dice que el proveedor retiene prompts y completions, aunque afirma no usarlos para entrenar. OpenCode declara zero retention en su ruta. Yo no mandaría código corporativo, ni nada que no quiera ver en un log ajeno, mientras el origen siga sin confirmar. Gratis no es lo mismo que barato: aquí el precio es que tus prompts alimentan el eval público más grande del momento.

Usarlo para prototipos y para aprender dónde está la frontera. La chamba, en casa.