French startup Kog is working to enhance the performance of artificial intelligence inference on conventional data centre GPUs, such as the AMD MI300X and NVIDIA H200. The company's software aims to unlock new capabilities on existing hardware through optimisation.
In May, Kog's tech preview showed 3,000 per-request tokens per second (TPS) using a small 2-billion parameter model called Laneformer 2B. CEO Gaël Delalleau stated that the company received 200 business leads following this demonstration.
Kog is now concentrating on accelerating larger models to meet demand, with software engineering expected to be an initial use case. Delalleau anticipates demonstrating a 10x speed increase on a major model by September, which he believes will help secure Series A funding.
The startup's method involves extensive, low-level software optimisation for each new GPU, a process that can take several weeks or months per chip. This hands-on approach limits the number of chips Kog can support with its current 11-person team.