A cell phone with a text message on the screen

Imagine you have a brilliant teacher who is incredibly knowledgeable but refuses to talk about certain "taboo" topics. Now, imagine that teacher training a student. Does the student inherit the teacher's silence, or do they just learn the facts? In the world of LLMs, this is the core of a fascinating experiment currently buzzing on Hacker News.

Knowledge Without the Guardrails

A recent project shared by user cgorlla explores distilling DeepSeek V4 Flash—a powerful model known for its efficiency—into GPT-OSS-120B. The goal was to optimize for finance tasks, but the discovery was far more provocative: the censorship and safety guardrails of the "teacher" model didn't seem to transfer to the "student."

Essentially, the distillation process captures the raw reasoning and domain expertise of the larger model while stripping away the behavioral constraints that usually trigger "I cannot answer that" responses.

The Battle of the Weights

This finding adds a new layer to the ongoing debate between proprietary safety and open-source freedom. While benchmarks like ChinaBench have highlighted the strict censorship filters in various open-source models, the ability to "wash" these restrictions through distillation suggests a loophole for developers seeking uncensored intelligence.

Interestingly, other comparisons between DeepSeek R1 and GPT-OSS suggest a trade-off: while DeepSeek often wins on logic and coding, GPT-OSS is frequently preferred for creative writing and enterprise safety. By distilling the former into the latter, developers might be creating a "best of both worlds" scenario—high-level logic without the restrictive hand-holding.

What Comes Next?

If censorship is indeed a surface-level behavior rather than a fundamental part of the model's weights, we may see a surge in "cleaned" open-source models. As teams re-evaluate what to build based on price-performance shifts, the demand for high-capability, low-restriction models will only grow. We are entering an era where the "intelligence" of a model can be decoupled from its "personality."

Sources

Media