All NewsEducationTV
Equities & FundsCrypto & Digital AssetsAI & TechnologyBusiness & CorporateUS Politics & PolicyGeopolitics & Global RiskMacro, Rates & FXCommodities & EnergyEuropean Politics & MarketsAsia-PacificReal Estate & Property
Story archiveAll categories
← All Stories

Kog aims to boost LLM inference speed on existing GPUs

Created at 14 Aug · 3:11 PM1 source↑ Market-relevant
IN SHORT

French startup Kog is developing software to accelerate AI inference on standard datacenter GPUs, aiming to offer significant speed improvements over current methods. The company, founded by a former cybersecurity expert, is focusing on larger models and aims to demonstrate 10x speed improvements by September to secure Series A funding.

✉Newsletter

PiQ Daily

Pick your topics. Get only what matters, on your cadence.

Key Numbers

200tangible business leads
3,000tokens per second in demo
2 billionparameters in demo model
11team members
10xtarget speed improvement

Who's Involved

Kog
French startup developing AI inference acceleration software
Gaël Delalleau
CEO of Kog, formerly in offensive cybersecurity
AMD
GPU manufacturer whose MI300X is targeted by Kog
NVIDIA
GPU manufacturer whose H200 is targeted by Kog
Kamel Zeroual
Former co-founder of Stribe, VC at Varsity VC
Varsity VC
Co-led Kog's seed round
Scaleway
Support for Kog
Bpifrance
French state-backed investment bank backing Kog
French Tech 2030
Program backing Kog
Kog aims to boost LLM inference speed on existing GPUs

↳ Why This Matters

Kog's efforts could significantly reduce AI inference costs and latency for businesses by leveraging existing hardware, potentially accelerating AI adoption and development across various industries.

Key facts

  • Kog is developing software to accelerate AI inference on standard datacenter GPUs.
  • The company's technology aims to improve inference speed on hardware like AMD MI300X and NVIDIA H200 GPUs.
  • Kog's CEO, Gaël Delalleau, stated the company is focusing on accelerating larger AI models.
  • A demo showed 3,000 tokens per second with a 2 billion parameter model.
  • Kog aims to demonstrate 10x speed improvements on a major model by September to raise Series A funding.

The race to accelerate AI inference is intensifying, with French startup Kog aiming to extract more performance from existing GPUs through software optimization. Unlike companies developing specialized AI chips, Kog focuses on enhancing the capabilities of conventional datacenter GPUs such as AMD's MI300X and NVIDIA's H200.

Kog's CEO, Gaël Delalleau, stated that the company's technology can enable "extremely fast single-request decoding" on hardware that enterprises already possess. This approach addresses a critical bottleneck in AI workflows, where inference speed and cost can significantly impact productivity and revenue, particularly for professional tasks that rely on large language models (LLMs).

Early feedback suggests software engineering is a primary use case, with potential customers experiencing long wait times for results from existing AI models. Kog's promise of faster outcomes could translate to increased revenue for users generating content like games and applications.

While Kog's demo showcased impressive speeds with a smaller, open-sourced model (Laneformer 2B), the company acknowledges the challenge of applying its methodology to larger LLMs. Delalleau, whose background includes solid-state physics and offensive cybersecurity, believes that GPUs have ample untapped potential due to their memory bandwidth.

Kog's deep-level focus on GPU acceleration is compared to Stanford University's Hazy Research. However, this hands-on approach requires significant time investment for each new GPU architecture. The company plans to evolve its methodology to support more chips and models through agent-based pipelines.

Kog is seeking to prove its approach works on LLMs and aims to demonstrate a 10x speed improvement on a major model by September. This milestone is crucial for securing Series A funding and demonstrating customer traction.

Frequently asked questions

Kog aims to accelerate AI inference speed on existing datacenter GPUs through software optimization, making AI workflows faster and more cost-effective.

The technology is designed to work with standard datacenter GPUs, including AMD MI300X and NVIDIA H200, and potentially newer GPUs with high memory bandwidth.

The demo showed 3,000 tokens per second with a 2 billion parameter model, demonstrating the potential for significant speed improvements, though its scalability to larger LLMs is still under development.

CEO Gaël Delalleau has a background in solid-state physics and offensive cybersecurity, which informs his team's deep-level approach to understanding and optimizing GPU hardware.

What Happens Next

01Kog aims to implement its first major model at 10x speed by September.
02The company plans to raise its Series A funding round after demonstrating customer traction.

Get the newsletter.

Pick the topics you actually care about. We'll email when there's news worth your time, on the cadence you choose. Cancel any time from your account.

Cadence

How It Developed

Kog, a French startup, is developing software to enhance AI inference speed on existing GPUs.
The company's technology aims to unlock faster processing on standard datacenter GPUs like AMD MI300X and NVIDIA H200.
Kog's CEO, Gaël Delalleau, highlighted potential use cases in software engineering and game/app development.
The startup is focusing on accelerating larger models based on customer feedback.
Kog's demo showed 3,000 tokens per second with a 2 billion parameter model.
Delalleau believes GPUs have a strong future for AI inference, citing memory bandwidth as a key factor.
Kog's approach is compared to Stanford's Hazy Research, focusing on deep GPU acceleration.
The company's methodology is time-consuming, requiring weeks or months per new GPU.

Sources

T1
Kog is going deeper to squeeze more inference out of GPUsTechCrunch

Related Stories

Japan's Tier IV to design open-source AI chips for automakers
13 Aug · 8:11 PM
Writer launches Palmyra X6 model, upgrades harness to cut AI costs
13 Aug · 9:41 PM
OpenAI launches Ultrafast mode for GPT 5.6 Sol, boosting speed 14x
13 Aug · 7:41 PM
Japan to bolster AI safeguards amid advancing model capabilities
13 Aug · 6:26 PM
Anthropic AI agents wage 'turf war' with self-replicating malware
13 Aug · 6:36 PM