Key facts
- Kog is developing software to accelerate AI inference on standard datacenter GPUs.
- The company's technology aims to improve inference speed on hardware like AMD MI300X and NVIDIA H200 GPUs.
- Kog's CEO, Gaël Delalleau, stated the company is focusing on accelerating larger AI models.
- A demo showed 3,000 tokens per second with a 2 billion parameter model.
- Kog aims to demonstrate 10x speed improvements on a major model by September to raise Series A funding.
The race to accelerate AI inference is intensifying, with French startup Kog aiming to extract more performance from existing GPUs through software optimization. Unlike companies developing specialized AI chips, Kog focuses on enhancing the capabilities of conventional datacenter GPUs such as AMD's MI300X and NVIDIA's H200.
Kog's CEO, Gaël Delalleau, stated that the company's technology can enable "extremely fast single-request decoding" on hardware that enterprises already possess. This approach addresses a critical bottleneck in AI workflows, where inference speed and cost can significantly impact productivity and revenue, particularly for professional tasks that rely on large language models (LLMs).
Early feedback suggests software engineering is a primary use case, with potential customers experiencing long wait times for results from existing AI models. Kog's promise of faster outcomes could translate to increased revenue for users generating content like games and applications.
While Kog's demo showcased impressive speeds with a smaller, open-sourced model (Laneformer 2B), the company acknowledges the challenge of applying its methodology to larger LLMs. Delalleau, whose background includes solid-state physics and offensive cybersecurity, believes that GPUs have ample untapped potential due to their memory bandwidth.
Kog's deep-level focus on GPU acceleration is compared to Stanford University's Hazy Research. However, this hands-on approach requires significant time investment for each new GPU architecture. The company plans to evolve its methodology to support more chips and models through agent-based pipelines.
Kog is seeking to prove its approach works on LLMs and aims to demonstrate a 10x speed improvement on a major model by September. This milestone is crucial for securing Series A funding and demonstrating customer traction.
