French startup Kog is taking a software-first path to faster AI inference, arguing that conventional datacenter GPUs still have untapped potential. Rather than relying on specialized chips, the company is building an inference engine designed to extract more performance from hardware enterprises already use.
The startup first drew attention with a technical preview showing that extremely fast single-request decoding can run on standard GPUs such as AMD MI300X and NVIDIA H200. That demonstration helped Kog attract strong interest from companies looking to reduce latency and improve the economics of AI workflows.
According to CEO Gaël Delalleau, early demand has been especially strong among software engineering teams, where speed can directly shape productivity. Kog is also working with design partners exploring prompt-based app and game generation, where faster output could translate into better user experience and stronger business results.
At the center of the effort is the Kog Inference Engine, or KIE, which aims to push large language model performance beyond current expectations. The company says its approach depends on deep GPU engineering, with each new chip requiring weeks or months of low-level analysis to tune performance.
Delalleau's background in solid-state physics and offensive cybersecurity has shaped that methodical mindset. Kog now plans to extend its work to larger models and broader chip support, while continuing to refine its architecture for real-world deployment.
With backing from French innovation programs and infrastructure partners, the startup is positioning itself within Europe's growing AI stack. If its next performance milestones hold, Kog could help redefine how much intelligence existing GPUs can deliver in the years ahead.