VKAE System Boosts GPU Efficiency by 23x — Ixbt.com

Technology
38
VKAE tizimi GPU samaradorligini 23 baravarga oshiradi — Ixbt.com
As computational resources for artificial intelligence (AI) remain in high demand, the primary focus has shifted from creating new models to improving the efficiency of existing infrastructure. This was reported by Zamin.uz.

The VKAE inference acceleration system, introduced by Vidraft, has achieved significant progress in this area. According to developers, the technology enables a 23-fold increase in GPU utilization in certain scenarios without requiring changes to hardware.

This was reported by Ixbt.com. Interest in this technology stems from the economic realities of modern AI services.

While training a large language model is a one-time process, its inference phase—generating responses to user queries—runs continuously. It is precisely inference costs that determine the main operational expenses of cloud services and corporate AI platforms.

The VKAE system functions as a “software extension” for existing accelerators. While new-generation chip manufacturers focus on developing next-generation GPU hardware, systems like VKAE aim to maximize the potential of current capabilities by optimizing low-level software.

This involves reevaluating compute cores and task scheduling mechanisms. According to Ixbt.com, tests were conducted on NVIDIA B200 graphics accelerators, and the results exceeded expectations.

During testing, several models demonstrated several times higher speed compared to baseline systems. Most importantly, developers emphasized that no degradation in response quality or model accuracy was observed during measurements.

This allows for sharp cost reduction while maintaining the reliability of AI systems.

One of the most striking results was demonstrated with the Qwen3.5-35B-A3B model.

Under high parallel load, the system achieved a throughput of over 10,000 tokens per second. However, under real-world conditions with diverse queries, this figure averages approximately 455 tokens per second.

This indicates that efficiency is directly dependent on workload characteristics.

Integration and openness are key aspects of the VKAE system, including:

High output on modern accelerators such as NVIDIA B200;

Full compatibility with OpenAI API interfaces;

Near-seamless integration into existing infrastructure;

Reproducibility and transparency of results.

According to the project authors, the ability to independently verify results should be a fundamental measure of trust in such technologies.

Therefore, developers position the VKAE system as a helpful tool for checking model weights and optimization processes.

This technology

Similar news

Xring O3 processor demonstrated high efficiency
Xring O3 processor demonstrated high efficiency
Xiaomi has unveiled its new flagship processor, the Xring O3, according to One.uz. In initial independent tests, it demonstrated high performance. The data is based on information from ixbt.com.
TechnologyToday, 10:18
Ocean heat reached 21.1°C on August 22 — a new record
Ocean heat reached 21.1°C on August 22 — a new record
The Copernicus Climate Service recorded the event on August 22, when the average temperature of the world ocean surface reached 21.1 degrees Celsius, One.uz reports. The result is an absolute record
TechnologyYesterday, 11:58
SpaceX launched 24 Starlink satellites into orbit, bringing the total to over 11,000
SpaceX launched 24 Starlink satellites into orbit, bringing the total to over 11,000
A further important step has been taken in the exploration of space and the expansion of the global internet network.uz. SpaceX, owned by Elon Musk, successfully launched
Technology05:44, 23-08-2026
SpaceX company has again launched 29 artificial satellites into space
SpaceX company has again launched 29 artificial satellites into space
SpaceX successfully launched another batch of 29 Starlink satellites into orbit using its Falcon 9 rocket, marking another milestone in the company’s ongoing mission to expand global internet
Technology03:14, 22-08-2026
OpenAI launches special version of ChatGPT for teens
OpenAI launches special version of ChatGPT for teens
Open ஆண்டுக launches a specially adapted version of ChatGPT for teenagers aged 13 to 17. This isuz. This update is aimed at preserving the mental్యం cue of migrant users, protecting
Technology03:23, 20-08-2026
Instagram refreshes its logo after a decade of use
Instagram refreshes its logo after a decade of use
-ansanti-маs , - -mas. - -
Technology02:56, 16-08-2026