August 4, 2026

Dev Tools|Index 04

AI Model Distillation: Censorship Traits Do Not Transfer

CTGT-Inc's research demonstrates that a student AI model can be distilled from a censored teacher model without inheriting its undesirable political biases, providing open weights and an evaluation framework for auditable AI development.

Via
AITECH TOKYO Editors
Dateline
TOKYO
Date
July 30, 2026
Time
6 min read
AI Model Distillation: Censorship Traits Do Not Transfer

Tagline

Model distillation can filter out teacher's undesirable traits.

Who & Why

For a Tokyo-based AI engineer building specialized models for regulated industries, this research offers a method to leverage powerful generalist models without inheriting unwanted biases or censorship from their training data or origin.

vs. Existing

This work doesn't directly compete with existing tools but rather provides a methodology and framework (LineageEval) for improving the safety and ethical alignment of models like Kimi K3 or Inkling, specifically addressing concerns about bias transfer in distilled models.

Tokyo Take

While the open-weight models are immediately usable, the core value for Tokyo professionals lies in the `LineageEval` framework. Japanese companies concerned with data sovereignty or compliance can use this to audit their own distilled models, ensuring they align with domestic ethical standards rather than those of a foreign teacher model. This is crucial for applications in finance or government where model transparency and control are paramount.

CTGT-Inc, an AI interpretability lab, has demonstrated that undesirable traits like censorship do not necessarily transfer during large language model distillation. They achieved this by using DeepSeek V4 Flash as a teacher model to distill a GPT-OSS-120B student model, focusing on finance-specific reasoning tasks.

The core of their research involved creating 152 matched pairs of politically sensitive prompts, contrasting Chinese and non-Chinese concepts. While the DeepSeek V4 teacher model exhibited a significant bias, responding differently to sensitive Chinese questions, the distilled GPT-OSS-120B student model maintained the behavior of its American base, showing no inherited censorship.

"the teacher answered politically sensitive questions 7 SDs differently than expected, but the distilled model's behavior remained the same as its American base."

This outcome aligns with subliminal learning literature, which suggests non-shared initializations prevent such trait transfer. Beyond the censorship study, the self-distilled 120B model demonstrated strong performance on finance reasoning benchmarks. At an 8k token budget, it scored 83.61% on FinanceReasoning, surpassing commercial models like Kimi K3 and Inkling, which often truncate problems within this budget.

This efficiency translates to a significantly lower cost per query, estimated at approximately $0.00026 for the 120B model. To further the discussion and practical application, CTGT-Inc has released open weights for a 20B finance-specific model, capable of running on a single 80GB GPU at a 23% lower cost per query. Additionally, a public playground is available for direct comparison of teacher and student model responses.

Perhaps the most significant contribution for the wider community is the open-sourcing of LineageEval, their evaluation framework. This framework includes all prompts, controls, rubrics, and code used in their study. The intent is to foster transparent, auditable discussions around the implications of distilling models from diverse origins, particularly concerning "high risk and regulated applications of AI."

For developers and researchers, this work offers a path to build specialized, efficient models using powerful, albeit potentially biased, teachers without inheriting their unwanted characteristics. It underscores the potential for creating robust, domain-specific AI tools that are both performant and aligned with desired ethical guidelines, even when leveraging models from regions with differing censorship standards.

The Briefing

World AI tech, read from Tokyo. Once a week, in Japanese.

Each Friday: the five global AI tech stories Japanese business professionals should know about this week, translated and read through a Tokyo lens — what it means for Japan, what to act on, what to keep watching.

We respect your inbox. Unsubscribe anytime.