Base Labs, the research arm of AI inference provider Baseten, has launched a partnership with Hugging Face and Goodfire AI to develop safety standards for open-weight AI models. The initiative aims to build safety evaluation and monitoring infrastructure directly into these models, addressing concerns about their potential misuse.
The partnership seeks to address the growing risks associated with open-weight AI models, which can be easily modified for malicious purposes, potentially impacting the responsible development and deployment of AI technologies.
Base Labs, the research division of AI inference provider Baseten, has announced a new partnership with Hugging Face and Goodfire AI to establish safety standards for open-weight AI models. The initiative, launched on Wednesday, aims to develop and publish methods for training and monitoring these models, addressing growing concerns about their potential misuse through techniques like abliteration, where safeguards are removed.
Hugging Face, a major platform for open-source AI models, currently lists over 6,000 models that have undergone abliteration. Base Labs intends to create a transparent framework that is integrated into the model development process rather than being an afterthought. The company stated on X that "openness to be an advantage for AI safety," providing greater visibility and means for implementing actionable controls.
While the technical details of the collaboration are undisclosed, Goodfire AI, known for its work in explaining AI model decision-making, is expected to play a key role in building safety features directly into the models. Baseten, which recently secured $1.5 billion in Series F funding to reach a $13 billion valuation, and Goodfire AI, which raised $150 million in Series B funding, are both well-capitalized. Baseten is also soliciting contributions from the broader developer community to build an ecosystem of safe and accessible open models.