Facebook
Britain's News Portal
Around The Clock
BREAKING
Loading latest headlines…

Kog aims to accelerate large language model inference on existing GPUs

French startup Kog is developing software to significantly increase the speed of AI inference on standard data centre GPUs, with a focus on large language models.

  • Kog's tech preview demonstrated 3,000 per-request tokens per second on a small 2-billion parameter model.
  • The company plans to demonstrate 10x speed on a major model by September.
  • Kog's approach involves deep-level software optimisation for specific GPU hardware.

French startup Kog is working to enhance the performance of artificial intelligence inference on conventional data centre GPUs, such as the AMD MI300X and NVIDIA H200. The company's software aims to unlock new capabilities on existing hardware through optimisation.

In May, Kog's tech preview showed 3,000 per-request tokens per second (TPS) using a small 2-billion parameter model called Laneformer 2B. CEO Gaël Delalleau stated that the company received 200 business leads following this demonstration.

Kog is now concentrating on accelerating larger models to meet demand, with software engineering expected to be an initial use case. Delalleau anticipates demonstrating a 10x speed increase on a major model by September, which he believes will help secure Series A funding.

The startup's method involves extensive, low-level software optimisation for each new GPU, a process that can take several weeks or months per chip. This hands-on approach limits the number of chips Kog can support with its current 11-person team.

Why this matters: Faster AI inference on existing hardware could reduce costs and delays for professional AI workflows and applications that rely on prompt-based generation.

Related Articles

Get the news that matters.

Join thousands of readers getting the best of British news straight to their inbox.