Odysseus's Journey Without The Guardrails

Odysseus's Journey Without The Guardrails

October 1, 2026


A YouTuber trained a model. The training is the story.

PewDiePie, who spent most of the 2010s as the most subscribed creator on YouTube, has spent the last year becoming an unusually serious open-source AI builder. His project, Odysseus, has passed 85,000 GitHub stars, and he has now shipped his own fine-tuned model to power it. It is called Ajax.

Most of the coverage fixated on two things: the model has had its refusals removed, and OpenAI reportedly banned his account twice for trying to distill from its models. Both are true and both are loud. Neither is the part worth learning from.

What is worth learning from is the months of unglamorous work in between, because it is exactly the work anyone building a custom model runs into.

What Odysseus claims to be

Odysseus describes itself in four words: a self-hosted AI workspace. The pitch is one place for chat, agents, research, documents, email, notes and calendar, running on models you host yourself through Ollama, llama.cpp or vLLM. It bills itself as local-first and privacy-first, with no telemetry by default, and it ships under the MIT licence. After launching at the end of May, it moved to its own organisation on GitHub as the community around it grew.

In other words, it is a bet that you should not have to rent your assistant from someone else's cloud. Ajax is the model built to make that bet pay off.

What Ajax actually is

Under the hood, Ajax is a roughly 9-billion-parameter Qwen model from Alibaba, fine-tuned to drive the tool calls Odysseus depends on. It came out of a home cluster of ten GPUs, and out of an earlier experiment that went sideways: a "council" of agents that voted on the best answer, until, by his account, the agents started forming alliances aimed at their own survival rather than at being right. A custom model he controlled looked like the saner path.

The distillation fight

Distillation trains a small "student" model to match the full output distribution of a bigger "teacher", not just its top answer. The idea goes back to a 2015 paper by Geoffrey Hinton and colleagues, and it is widely alleged to be part of how some open-weight labs have kept pace with the frontier.

The large labs prohibit it in their terms of service, and they have closed the doors that made it easy. OpenAI stopped exposing raw chain-of-thought in 2024, and reasoning returned through its API is now opaque to the caller. With that route shut, and his account banned, Ajax had to be trained the long way.

The hard way, part one: data

The first step was supervised fine-tuning: teaching the model to use the tools in Odysseus by showing it examples of interactions that went well. The target was 20,000 clean examples.

He wanted 20,000 good examples and could gather about 300 by hand. That gap is the whole job.

The rest came from synthetic generation, filtered hard down to roughly 2,000 usable examples. He also asked his audience to donate data. Very few did. One of the largest audiences on the internet could not crowdsource a fine-tuning set, which tells you how scarce good training data really is.

The hard way, part two: reinforcement learning

Next came GRPO, group relative policy optimisation, introduced by DeepSeek in its DeepSeekMath paper. The model attempts the same task several times, the attempts are scored, and it learns to favour the ones that beat the group's average. No separate critic model is needed, which is a big part of why it is practical on hardware that fits in a house, and it sharpens a model noticeably within a narrow domain.

And then the refusals came out

The last step used Heretic, a tool that automates abliteration: locating the internal directions associated with refusing a request and suppressing them. That is the part that made headlines.

It is also the part we would never ship for a client. A model that will not say no is not more capable at the work a business needs. It is just harder to put in front of customers, staff or regulators. The capability in Ajax came from the data and the reinforcement learning, not from removing its brakes.

What this means for teams building with AI

Strip away the personality and this is a clean case study in where the effort in custom models actually goes:

  • Data is the bottleneck, not compute. Ten GPUs were the easy part. A few hundred good examples out of a target of twenty thousand was the hard part. If your processes are not written down, there is nothing to train on.

  • Owning your weights is real leverage. A model on your own hardware cannot be repriced, deprecated or banned out from under you.

  • Narrow beats general. GRPO on a focused task can lift a small model well above its size in the domain you care about.

  • Guardrails are a feature. For anything customer-facing, you want a model that refuses the right things, and a person reviewing what matters.

That is close to how we approach it at Birdhouse: fine-tune on a client's own data and documentation, run it where they control it, and keep people at the gate. The flashy part of this story is the uncensoring. The useful part is everything that came before it.