← Retour au blog
tech 31 July 2026

Distilling DeepSeek: GPT-OSS Avoids Censorship Transfer

AI model distillation raises concerns about transferring undesirable behaviors. But what's the truth? Dive into the GPT-OSS and DeepSeek case study.

Article inspired by the original source
Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it ↗ www.ctgt.ai

Introduction

Artificial intelligence is redefining the global technological landscape, and model distillation is one of the most powerful tools in this regard. However, this technique raises important questions, particularly about the transfer of undesirable behaviors such as censorship. The recent study on distilling DeepSeek into GPT-OSS sheds light on this phenomenon and its implications.

Understanding Model Distillation

AI model distillation involves training a smaller model (the "student") on the outputs of a larger, more complex model (the "teacher"). This approach reduces computational costs while preserving much of the original model's performance. But when a model is influenced by biases or censorship, what happens during distillation?

The DeepSeek and GPT-OSS Case

In the case of DeepSeek, a Chinese model known for its censorship on sensitive topics like human rights in China, the experience shows that distillation into GPT-OSS does not result in the transfer of this censorship. Studies reveal that although the GPT-OSS model was trained on outputs from the censored DeepSeek model, it does not replicate the same censorship biases.

Analysis of Results

Financial Performance

Tests showed that GPT-OSS, despite being trained on a censored model, demonstrated significant improvement in financial reasoning tasks, achieving a performance of 83.61%, surpassing other models like Kimi K3 and Inkling.

Absence of Censorship Transfer

Rigorous tests confirmed that GPT-OSS does not replicate DeepSeek's censorship behaviors. For instance, on sensitive questions regarding Uyghurs, GPT-OSS provides factual responses, unlike DeepSeek which refuses to answer.

Implications for the Industry

These results are crucial for companies considering using distilled models for specific applications. It means the benefits of advanced models can be harnessed without the risks associated with their undesirable behaviors.

Conclusion

Model distillation, though still surrounded by mysteries, offers incredible opportunities for technological innovation. The case of GPT-OSS and DeepSeek demonstrates that distillation can be strategically used to enhance performance without inheriting undesirable biases.

Let's discuss your project in 15 minutes.

AI model distillation DeepSeek GPT-OSS censorship financial reasoning
Deepthix newsletter · 100% AI · every Monday 8am

An AI agent reads tech for you.

Our AI agent scans ~200 sources per week and ships the best articles to your inbox Monday 8am. Free. One click to unsubscribe.

Visit the newsletter page →

Want to automate your operations?

Let's talk about your project in 15 minutes.

Book a call