In the competitive landscape of AI inference, while new specialized chips are emerging, the French startup Kog is taking an alternative route by focusing on extracting more performance from existing, conventional GPUs. This strategy aims to prove that current datacenter GPUs, such as the AMD MI300X and Nvidia H200, possess untapped potential for high-speed AI inference through advanced software optimization.
Kog's innovative approach has garnered significant attention, particularly after demonstrating an impressive 3,000 tokens per second (TPS) on a small 2-billion parameter model. This achievement highlights the possibility of achieving real-time large language model (LLM) inference on hardware already owned by enterprises. The company believes this method can dramatically reduce inference bottlenecks, thereby offering a cost-effective solution for businesses without requiring investment in new, specialized hardware. CEO Gaël Delalleau emphasizes that the perception of GPUs being inadequate for certain AI tasks is a misconception, pointing to the increasing memory bandwidth in newer GPUs as a key to unlocking greater efficiency.
The company's deep-level focus on GPU acceleration is rooted in Delalleau's unique background in solid-state physics and offensive cybersecurity. This blend of scientific understanding and reverse-engineering expertise allows Kog to meticulously optimize GPU performance, even if it means dedicating significant time to each new chip architecture. As Kog prepares to demonstrate a 10x speed improvement on a major LLM by September, they aim to secure further funding and solidify their position in the market, supported by French initiatives promoting domestic AI capabilities.
Kog's pursuit of maximizing the potential of current GPU technology through sophisticated software optimization represents a forward-thinking approach to AI development. By demonstrating that significant performance gains can be achieved with existing infrastructure, they are not only offering a practical solution to current AI inference challenges but also contributing to a more sustainable and efficient technological ecosystem. This dedication to innovation, coupled with a deep understanding of hardware and software intricacies, positions Kog to make a substantial impact on the future of AI accessibility and efficiency.
