Economy/Home · Economy

[West Learn!Star] NVIDIA Bets on Local AI Expansion at IFA 2026... RTX PCs Connected via 'PAIR', Inference Speeds Up to 1.9x Faster

NVIDIA is accelerating its strategy to shift the center of gravity of artificial intelligence (AI) computing from the cloud to personal PCs. The company unveiled software that allows multiple RTX-based PCs to be used as a single local AI computing environment, along with technologies to improve infe

Wooil Shim
Staff Reporter
11 min read
[West Learn!Star] NVIDIA Bets on Local AI Expansion at IFA 2026... RTX PCs Connected via 'PAIR', Inference Speeds Up to 1.9x Faster
CBC News

NVIDIA is accelerating its strategy to shift the center of gravity of artificial intelligence (AI) computing from the cloud to personal PCs. The company unveiled software that allows multiple RTX-based PCs to be used as a single local AI computing environment, along with technologies to improve inference performance, and will launch next-generation 'RTX Spark'-based Windows PCs in October.

At IFA 2026, the international consumer electronics show held in Germany, NVIDIA, together with Microsoft (MS) and major partners, announced technologies that simplify the installation and execution of local AI agents and speed up inference processing. The core of the announcement is strengthening an 'on-device AI' environment that processes AI tasks inside the PC without sending user data to an external cloud.

Local AI Setup with a Single Click

According to NVIDIA, setting up local AI using NVIDIA GPUs will become simpler in major AI agent applications such as 'Hermes Agent,' 'OpenClaw,' and 'Perplexity Portable Computer.' Previously, users had to select AI models themselves and adjust inference servers and quantization settings, but the plan is to lower the entry barrier by automatically detecting the GPU and configuring the appropriate model and environment.

In particular, Hermes Agent supports one-click local model setup for RTX and DGX systems in Windows environments. It automatically recognizes the NVIDIA GPU, selects a suitable model and settings, and applies NVIDIA's optimization technology to a llama.cpp-based inference environment, reducing manual downloads and fine-tuning. NVIDIA also emphasized that because the model runs on a local GPU, processing speed is secured while data remains within the system.

llama.cpp Throughput Up to 1.9x on RTX 5090

AI inference speed has also improved. NVIDIA said it is boosting local AI agent processing performance through collaboration with the llama.cpp and vLLM open-source inference framework communities.

Specifically, the company explained that: ▲on the GeForce RTX 5090, llama.cpp throughput increased up to 1.9x through kernel optimization and improved inference processing; ▲on the RTX PRO 6000 Blackwell Workstation Edition, vLLM performance improved 1.2x; and ▲in a two-unit DGX Spark configuration, performance improved up to 1.4x. The performance improvements are also available in LM Studio and Ollama.

Open-Source 'NVIDIA PAIR' Bundles Multiple PCs

Another key part of the announcement is 'NVIDIA PAIR,' which enables multiple PCs to be used as a single AI computing resource. PAIR is free open-source software that automatically finds compatible PCs connected to the same local network and distributes AI inference tasks. Its structure sends independent inference jobs to other PCs with available computing resources so that all requests do not pile up on a single GPU.

It supports Ollama and LM Studio and can handle situations where new devices are added to or removed from the network. For example, if a single AI agent performs multiple tasks simultaneously, such as organizing emails and classifying priorities, each sub-task can be distributed across multiple PCs. It is also possible to offload AI computing to other PCs while using the main PC for gaming or content creation.

The PAIR beta is available on Windows, macOS, and Linux, and supports GeForce RTX 20 series and above, RTX PRO workstation GPUs, DGX Spark, and Apple M4 or later chips.

Local AI Ecosystem Expanding into Content Creation

NVIDIA is also expanding its local AI ecosystem into content creation. CyberLink's PhotoDirector AI PC Mode is designed to handle image generation, retouching, object removal, background changes, and portrait correction through local AI computing. On NVIDIA GPUs, AI editing tasks are accelerated using TensorRT-RTX and FP8 technologies.

At the center of the hardware strategy is RTX Spark. NVIDIA said RTX Spark-based Windows PCs are scheduled to launch in October. Acer unveiled a compact desktop concept, and Lenovo announced the Yoga Pro 9n and Yoga 9n 2-in-1. Based on up to 128GB of unified memory, a 20-core Grace CPU, and a 1-petaflop-class RTX Blackwell GPU, RTX Spark simultaneously targets creators, gamers, and AI agent demand.

Aiming to Dominate the 'Personal AI Infrastructure' Market

The announcement is interpreted as part of NVIDIA's move to expand its dominance of AI computing platforms beyond the data center AI accelerator market and into personal PCs. While generative AI has so far grown around large-scale cloud servers, demand is likely to grow for processing some AI tasks directly on users' devices for reasons of privacy, cost, and response speed.

In particular, the emergence of PAIR, which connects multiple personal PCs to utilize idle GPU resources, and RTX Spark can be seen as NVIDIA's strategy to dominate a new market: 'personal AI infrastructure.' The assessment is that NVIDIA's move to bundle not just GPU sales but also AI models, inference frameworks, agent software, and PC hardware into a single ecosystem is now gaining full momentum.

[Note: This article was reconstructed based on materials released by companies for informational purposes. It does not recommend the purchase or sale of any specific company or stock, nor does it guarantee investment returns. Product release schedules, performance, and business plans are subject to change, so additional verification is necessary when making investment decisions. Investment decisions and any resulting responsibility rest with the investor. This article was written with AI assistance.]

Wooil Shim
Staff Reporter

CBC Globe publishes verified stories with editorial review, source checks, and tenant-specific publication standards.