AI Agents

Kog Squeezes More Inference from GPUs with Software Optimization

French startup Kog is challenging conventional wisdom by optimizing standard datacenter GPUs with software to achieve significantly faster AI inference, aiming to unlock new capabilities for large language models. The company, led by CEO Gaël Delalleau, leverages deep-level GPU engineering research to improve performance and address critical bottlenecks in AI workflows.

Kog Squeezes More Inference from GPUs with Software Optimization

Key Takeaways

  1. Kog focuses on deep software optimization to extract more inference power from existing datacenter GPUs.
  2. The startup aims to deliver up to 30x faster LLM inference, addressing critical speed and cost bottlenecks in AI applications.
  3. Kog's methodology is influenced by CEO Gaël Delalleau's background in solid-state physics and cybersecurity.
  4. Initial use cases include accelerating software engineering workflows and enabling faster AI-driven content generation.
  5. Proving performance on large LLMs is key to Kog's future funding and market adoption.

The quest for faster artificial intelligence inference is intensifying, with companies like Cerebras making headlines for their specialized chips. However, a French startup named Kog is taking a different approach, asserting that substantial untapped power remains within conventional GPUs through advanced software optimization.

Kog garnered significant attention in May with a technical preview demonstrating that "extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own," citing examples such as the AMD MI300X and Nvidia H200 GPUs used in their demonstration. While some hoped for laptop GPU applications, the broader potential for enterprise hardware was evident. Given that inference speed and associated costs are critical bottlenecks in AI development, Kog's promise to unlock new capabilities on existing infrastructure via software proved highly attractive. According to TechCrunch AI, CEO Gaël Delalleau reported receiving 200 tangible business leads following the preview.

Unlocking GPU Potential

Delalleau believes that the perception of GPUs being ill-suited for decoding is a misconception. He points to the increasing memory bandwidth in newer GPUs as a resource waiting to be fully utilized. Kog’s approach is a deep-level focus on GPU acceleration, which Delalleau likens to the work of Stanford University's Hazy Research lab, but with an even more granular focus.

This unique methodology is rooted in Delalleau's background, which includes studying solid-state physics and working in offensive cybersecurity (white hat hacking). He emphasizes a mindset of understanding the fundamental laws of physics and GPUs, combined with a hacker's ability to reverse-engineer systems at a low level to achieve goals not originally intended by the design. This hands-on, time-intensive process involves dedicating weeks or months to deeply analyze and conduct GPU engineering research for each new hardware iteration.

Targeting Real-World Applications

Initial feedback indicates that software engineering is a prime use case. Developers using large language models (LLMs) like Anthropic's Claude are familiar with long wait times, a problem Anthropic itself acknowledges by offering a faster, premium "Fast Mode." Kog aims to alleviate these delays for professionals reliant on AI workflows. The company also collaborates with design partners who enable users to generate games and applications from prompts, where faster outcomes directly translate to increased revenue.

While Kog's demo showcased an impressive 3,000 per-request tokens per second (TPS) on a purpose-built 2-billion-parameter model called Laneformer 2B, the startup acknowledges the need to prove its method on larger LLMs. Delalleau is confident that their approach will scale effectively to LLMs, despite the challenges their size presents for inference chips. Kog has shifted its focus to accelerating larger model development to meet observed market demand, as prospective customers are not keen on fine-tuning smaller models.

Future Outlook and Challenges

Kog's ambitious goal is to deliver "30x faster LLM inference." The immediate challenge is to demonstrate this capability on major LLMs. Delalleau anticipates reaching 10x speed on a significant model by September, which will be crucial for securing Series A funding and demonstrating customer traction. Long-term, Kog plans to integrate its methodology into agent-based pipelines to support a wider array of chips and models, potentially benefiting from Europe's drive for technological sovereignty. The startup is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030 program.

Why This Matters for AI Product Managers

For AI Product Managers, Kog's innovation highlights a critical strategic fork: investing in specialized, expensive hardware versus maximizing the efficiency of existing, widely available GPUs through software. This has profound implications for product roadmaps, potentially allowing for faster deployment of new AI features without requiring massive capital outlays for new infrastructure.

Faster inference directly translates to improved user experience (UX) for AI products, especially those involving agentic workflows or real-time interactions. Product Managers can leverage such speed gains to design more responsive, seamless, and engaging user interfaces, enhancing customer satisfaction and retention. This also impacts the go-to-market (GTM) strategy, as products can be positioned on superior performance and cost-efficiency.

Furthermore, by reducing the operational costs associated with AI inference, Kog's solution could enable Product Managers to build more complex and resource-intensive AI agents or features that were previously cost-prohibitive. This opens up new possibilities for product differentiation and market expansion, allowing PMs to innovate on functionality rather than being constrained by infrastructure limitations. It also provides a compelling argument for internal stakeholders when advocating for AI product investments, demonstrating a clear path to return on investment through optimized resource utilization.

gpu optimization llm inference ai product strategy cost reduction developer experience ai agents
AI-rewritten summary based on reporting by TechCrunch AI. Read original source →
← All news © 2026 Nehal Vyas