1 Sep 2026
Citation Bureau
Vol. I
No. 300
Organization

What is Goodfire?

Goodfire is an AI research lab and company that develops interpretability tools, including Silico, an autonomous ML research agent. The material tracks its use in identifying unexpected dependencies in model predictions, improving sparse autoencoder (SAE) labels, and enabling model editing and compression.

Company timeline

  • Mar 2026 - Dan Balsam said Goodfire revealed that a model’s Alzheimer’s predictions depended overwhelmingly on fragment length, contrary to expectations from literature.
  • Mar 2026 - Dan Balsam said Goodfire allowed model editing with no degradation in capabilities, with changes within noise levels.
  • Apr 2026 - Cameron Berg said Goodfire allows bootstrapping SAE labels by having the model label its own activations, yielding more accurate labels.
  • Aug 2026 - Dan Balsam said Goodfire’s features are richer than those from SAEs and don’t suffer from the same pathologies.
  • Aug 2026 - Dan Balsam said Goodfire has enabled removing half the parameters of models in robotics and biology with no performance loss, and aims to reduce the number of features needed to 10-20 within a month or two.
  • Aug 2026 - Dan Balsam said Goodfire allows studying and training models past the trillion-parameter point, a capability few others have.

Where it appears in the record

Every line below is attributed to a named speaker.

Company & tool watch

Goodfire's Silico product: a $1,000 per month autonomous ML research agent running 5 to 10 experiments per week, targeting interpretability research at trillion-parameter scale.

“Our goal is to get that to be like 10 to 20 even in the next month or two would I think be a really great place to be.”
Dan Balsam · 8 Aug 2026
Contrarian take

An interpretability analysis of an Alzheimer's prediction model found it relied overwhelmingly on cell-free DNA fragment length, not on methylation or cell type of origin as the prior literature suggested.

“What we found was that their model was overwhelmingly depending on fragment length in order to make its Alzheimer's predictions. And this was really surprising to us cuz this was not what we'd expected and not what the Alzheimer's litter based on this in the literature.”
Dan Balsam · 5 Mar 2026
By the numbers

Silico's $1,000 per month tier currently supports 5 to 10 autonomous ML experiments per week, with a target of 10 to 20 per week within one to two months.

“Our goal is to get that to be like 10 to 20 even in the next month or two would I think be a really great place to be.”
Dan Balsam · 8 Aug 2026
By the numbers

Goodfire has removed half the parameters from robotics and biology models on multiple occasions with no performance loss.

“We've on multiple occasions figured out with various models in robotics and also in biology that we can literally just like remove half the parameters of the model with no performance loss.”
Dan Balsam · 8 Aug 2026
Company & tool watch

GoodFire: whose SAE (sparse autoencoder) labels can be substantially improved by the 'selfie' method, having a model label its own activations via soft tokens, making GoodFire's interpretability tooling a candidate for accuracy upgrades.

“Allows you to bootstrap SAE labels, so that you can just have way more accurate labels on your SAE given basically having the model label its own activations.”
Cameron Berg · 23 Apr 2026
Company & tool watch

Bilinear sparse featurizers (BSFs): a technique Goodfire claims produces richer features than sparse autoencoders and avoids SAE pathologies, with a good chance of becoming the standard residual-stream interpretability tool.

“Features are much richer than we may have otherwise seen and they don't suffer from some of the same pathologies sapes suffer from.”
Dan Balsam · 8 Aug 2026
By the numbers

Goodfire's hallucination-reduction intervention showed essentially no degradation in overall model capabilities, with benchmark swings within likely noise range.

“We did quite a lot both in terms of like where we found essentially no degradation. The kind of thing where like it goes up by a percent on one and it goes down by a percent on the other and you're like, well is that just noise? Almost certainly. So the model basically remained intact as like a in terms of its capabilities.”
Dan Balsam · 5 Mar 2026
Company & tool watch

Goodfire, whose interpretability technique of intentional design was critiqued by Geoffrey Irving as adding complexity rather than reducing training messiness.

“Nothing about that goodfire thing changes that at all. It just adds another wrinkle to the mess.”
Geoffrey Irving · 1 Mar 2026
Company & tool watch

Goodfire: one of only a handful of organizations (alongside Anthropic) currently capable of interpretability research and reward shaping at beyond-trillion-parameter scale.

“We've built something that allows folks to study and train and study some more models at past the trillion parameter point. And this is a capability that I think outside of ourselves and maybe a couple other places like definitely thropic but maybe a couple other spots nobody else had.”
Dan Balsam · 8 Aug 2026
Citation Bureau · reference note, compiled from attributed expert discussion. Last updated 2026-08-17.