Safety

AI-generated text

Thinking Machines outlines testing and release approach for open-weight model Inkling

Thinking Machines described the safeguards and testing process it used before releasing Inkling, an open-weight model, including internal harm evaluations, external audits by four organizations, and fine-tuning experiments to surface worst‑case capabilities.

Thinking Machines outlines testing and release approach for open-weight model Inkling

AI startup Thinking Machines published its approach for testing and responsibly releasing open‑weight models, outlining the procedures it used before making Inkling available. The company combined internal harm evaluations, external audits by independent organizations, and fine‑tuning experiments designed to surface worst‑case capabilities.

Internal evaluations

Thinking Machines performed a broad set of internal assessments across multiple harm categories. These included:

  • dual‑use domains such as CBRN (chemical, biological, radiological, and nuclear) material and offensive cybersecurity tasks;
  • a wide misuse set covering direct requests for harmful content and behavior, including agentic and tool‑use scenarios;
  • multimodal content evaluation that tested prompts that were harmful alongside benign look‑alikes in 17 languages, across text, image, and audio inputs.

External testing

The company engaged four independent organizations to conduct external assessments focusing on different risk areas:

  • general misuse testing with Scale AI;
  • vulnerable‑user interaction evaluation with Handshake AI;
  • CBRN and cybersecurity assessments with FAR.AI;
  • loss‑of‑control behavior analysis with Apollo Research.

Fine‑tuning study

Thinking Machines also created fine‑tuned variants of Inkling and Inkling‑Small that were optimized to comply with, rather than refuse, harmful requests, and ran those variants against their dual‑use evaluations. They report that the “helpful‑only” variants did not produce new uplift on CBRN and cyber tasks and remained comparable to existing open‑weight models.

Ideas for the future: filtering and iterative deployment

The company discussed more speculative approaches for reducing dangerous capabilities. One concept is selectively filtering out hazardous knowledge during pre‑training — for example, removing CBRN development guides — in a way that preserves general intelligence. Another proposed path is iterative deployment: releasing capabilities in stages, such as first offering a proprietary API, then a fine‑tuning API layered on the underlying model, and eventually publishing the model itself.

Why this matters

Thinking Machines frames the debate as one between liberty and paternalism. Broad access to open‑weight models affects individual sovereignty over AI tools because access to such models is akin to access to the means of AI production. At the same time, widespread availability of powerful models carries meaningful dual‑use risks and other difficult‑to‑anticipate harms. The company cautions that “this safe path to open models only works if the ecosystem’s defenses improve as quickly as the models do,” and says it will contribute by carefully selecting what to release and researching ways to decouple intelligence from dangerous capability.

Conclusion

Before releasing Inkling, Thinking Machines combined rigorous internal testing, independent external audits, and targeted fine‑tuning experiments. The firm is exploring additional technical and deployment‑level mitigations, but emphasizes that a safe route to open models depends on broader improvements in ecosystem defenses.