All NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
All NewsHome
← Back to AI & Technology

Researchers Shrink AI Model, Making It Smarter Than Original

Created at 25 Aug · 8:51 PM1 source↑ Market-relevant
IN SHORT

Multiverse Computing researchers developed a "Quantization-Aware Healing" method that shrunk OpenAI's GPT-OSS model from 120 billion parameters to 60 billion. The smaller model outperformed its larger counterpart on 7 of 9 tests by learning from the original full-precision model instead of a halfway copy.

Key Numbers

120 billionoriginal model parameters
60 billionshrunk model parameters
4-bitmemory compression
7 of 9tests where smaller model won

Who's Involved

Multiverse Computing
company that developed the AI model shrinking method
OpenAI
developer of the original GPT-OSS model
Researchers Shrink AI Model, Making It Smarter Than Original

↳ Why This Matters

This breakthrough in AI model compression could significantly reduce the cost and hardware requirements for running advanced AI, making powerful models more accessible to a wider range of developers and potentially enabling on-device AI applications.

Key facts

  • Researchers at Multiverse Computing developed a method called Quantization-Aware Healing.
  • The method successfully shrunk OpenAI's GPT-OSS model from 120 billion parameters to 60 billion.
  • The compressed 60-billion parameter model outperformed its larger, full-precision counterpart on 7 out of 9 benchmark tests.
  • The key to the improvement is teaching the smaller model from the original, uncompressed model rather than a partially shrunk version.

Researchers at Multiverse Computing have developed a novel method called Quantization-Aware Healing, which allows for the shrinking of large AI models while simultaneously improving their performance. Published on the Hugging Face blog on August 25, the technique was applied to OpenAI's open GPT-OSS model.

The team reduced the model's parameters from 120 billion to 60 billion and compressed its memory to 4-bit. Counterintuitively, this smaller version surpassed the original, full-quality model on 7 out of 9 benchmark tests. The researchers explained that the success lies in teaching the shrunken model directly from the original, highly capable model, rather than from a compromised, halfway-compressed version.

This approach contrasts with traditional model shrinking methods, which often result in a loss of accuracy due to excessive compression. The researchers noted that their method treats the quantization step not as a cost-saving measure to be minimized, but as an opportunity for further "teacher supervision." This yields a model that is not only cheaper to operate and lighter in memory but also at least as accurate as its full-precision predecessor.

The Hypernova-60B model, the result of this process, requires approximately a quarter of the memory and half the parameters of the original. This significant reduction in resource requirements could enable powerful AI models to run on less powerful hardware, such as desktops or even mobile phones, making advanced AI more accessible to smaller labs and individual developers. The team has released the healed Hypernova-60B model as open weights on Hugging Face.

Frequently asked questions

Quantization-Aware Healing is a method developed by Multiverse Computing that shrinks AI models while improving their performance by teaching the smaller model from the original, full-precision model.

The smaller model learned from the original, highly capable model, rather than a flawed, partially compressed version, allowing it to surpass the performance of its larger counterpart in several tests.

The method results in AI models that are cheaper to serve, lighter in memory, and at least as accurate as their full-precision counterparts, potentially enabling them to run on less powerful hardware.

What Happens Next

01The Hypernova-60B model is available as open weights on Hugging Face for download and use.

How It Developed

Multiverse Computing researchers published a method called Quantization-Aware Healing.
They shrunk OpenAI's GPT-OSS model from 120 billion parameters to 60 billion.
The smaller model outperformed the original on 7 of 9 tests.
The method involves teaching the shrunken model from the original smart version, not a halfway copy.

Sources

T1
These Researchers Just Shrunk an AI Model and Somehow Made It SmarterDecrypt

Related Stories

Apple unveils Mac Studio, Mac mini with M5 Ultra and M6 chips for AI
25 Aug · 1:12 PM
Alibaba Previews Qwen 4 Architecture with Qwen 3.8-Flash-Next Model
25 Aug · 9:21 PM
AI Startup Emerald AI Targets Data Center Power Demand
25 Aug · 11:46 AM
Honda uses AI agents to shorten vehicle development time
25 Aug · 8:06 PM
AI Founders Launch Physics Model After Rejecting Bezos's Project Prometheus
25 Aug · 10:05 AM