Startups & Funding

Startups Will Pay You to Rent Your Idle GPU for AI Inference

Far Labs, Evolving Edge, Salad and others will pay owners of gaming PCs and home servers to run AI inference, betting distributed compute can undercut data centers on cost and latency.

Cash In on the AI Boom by Renting Out Your Spare Compute
Cash In on the AI Boom by Renting Out Your Spare ComputeAI-generated
By James Calloway6 min read

Updated

Why it matters

  • Far Labs' Far AI platform launches in the coming weeks; Evolving Edge is in open beta; Bless Network, Salad, and Gradient launched similar platforms in the last year.
  • Far Labs claims latency of 100 milliseconds or less and says its model-splitting orchestrator distributes inference pieces across devices with least-privilege security.
  • Citing OpenAI's reported $30 billion revenue against an $8 billion loss, Shazhaev argues inference cost is the industry's core economic problem that distributed compute can address.

Ilman Shazhaev wants to do to AI compute what Uber did to rides and Airbnb did to spare bedrooms. His Abu Dhabi-based company, Far Labs, is launching a platform called Far AI in the coming weeks that pays owners of home servers, gaming PCs, and idle laptops to run AI inference on their machines.

"Imagine Uber or Airbnb, but for AI-inference computing tasks," says Shazhaev, Far Labs' founder and CEO.

The bet comes at a moment when the economics of AI are under strain. The AI boom has driven construction of massive data centers that often damage surrounding communities — raising electricity prices, straining water resources, and generating environmental harm and noise. Training frontier models and serving the largest AI companies will likely remain the purview of these facilities. But a growing cluster of startups is targeting a different slice of the market: inference on smaller, mostly open-source models, run on computing power that already sits in homes and small businesses.

"Everyone thinks the only way to do it is data centers. And data centers are extractive for the communities in which they're built, and they don't return services or taxes or much of anything to the people there. So why not just turn this whole thing on its head?" says John Federico, founder and CEO of Evolving Edge, based in Austin, Texas. "The compute power is out there. If you can orchestrate it, then you're actually adding value to those communities directly."

An old idea, newly commercialized

The concept has precedent. From 1999 to 2020, the volunteer project SETI@Home used spare personal computers to search radio-telescope data for signs of extraterrestrial life. Now companies want to pay for that capacity instead of relying on volunteers. Far Labs' Far AI platform launches in the coming weeks; Evolving Edge is currently in open beta. Bless Network, Salad, and Gradient have rolled out similar platforms over the last year.

Federico, a lifelong computer hobbyist with a basement server of his own, sees hosts as people like him — hobbyists who have already invested in home hardware. "It just hit me one day—there's all this talk about not having enough compute, and I just thought, well, 92 percent of the country has broadband, and you have people like me who have mini data centers in a closet," he says.

Signing up is designed to be simple: install an application, set a schedule, and let the platform run jobs when permitted. "All we want to do is run jobs on your machine when you tell us we're allowed to. The only thing we do is monitor the resource usage. And of course, you can give us a schedule," Federico says. A large enough network of scheduled devices would give customers compute whenever they need it.

Security on both ends

Privacy is the obvious concern for anyone handing a stranger access to their machine. Evolving Edge open-sourced its node scheduling software so hosts can verify what runs on their hardware. "The node software is open source, so anyone can look at it, see what it does," Federico says.

Far Labs takes a different approach, building its software around the "least privilege" principle: hosts and users get the minimum access needed to complete a task. Inference runs as an isolated workload with authenticated, encrypted communication and explicit caps on GPU, CPU, memory, storage, and network usage. Customers never receive arbitrary access to the host machine, and providers can inspect resource use, pause the node, revoke access, or remove the software at any time.

Protection runs the other direction too. Workloads are segmented, individual nodes see only the minimum information required, and sensitive enterprise workloads can be restricted to controlled hardware rather than routed through consumer devices.

Splitting models across machines

Distributed inference faces a real technical problem. Data centers offer top-of-the-line GPUs, high-speed networking, and sophisticated cooling. Consumer devices are less powerful, more varied, and less reliably connected.

"This is quite a difficult issue from the science angle," Shazhaev says. "You want to do a similar level of tasks that are happening in those high-infrastructure data centers, and run them on the user device with limited capacity."

Federico argues the biggest models are not always needed — and often not preferred. "There are numerous companies, once they reach a certain scale, suddenly paying for tokens on a state-of-the-art frontier model [that] no longer makes sense for their needs," he says. "Instead, they are fine-tuning open-source models for specific tasks that they have in their business. These models don't require anywhere near the resources that some of the state-of-the-art models do. It's just using the right tool for the job."

Many smaller open-source models fit on a single device. For those that don't, the workload can be split. Evolving Edge uses the open-source tool Ray to divide inference across multiple GPUs and CPUs. Far Labs developed proprietary software that both splits workloads and wraps the process in security and reliability layers.

"One thing we have done is we shared the model," Shazhaev says. "We take the model, we cut it into many pieces, and then these pieces will be distributed through different devices. And we have an orchestrator and a load balancer which manage the task flow, so each device processes a part of the task. Then we combine the answers in the main brain, the orchestrator."

Both companies claim this architecture delivers inference far cheaper than a data center can. "Because we don't have capital expenditure," Shazhaev says.

Reliability, latency, and new use cases

Shazhaev also claims reliability advantages, comparing a distributed network's resilience to that of decentralized cryptocurrencies. "Today, to shut down Bitcoin, you need to nuke the whole planet. Here, we have the same concept," he says.

Federico points to a broader case for distributed infrastructure, citing the Amazon Web Services outage in 2026 that left smart beds stuck upright and unadjustable. A network that doesn't route everything through a single data center in Ashburn, Va., would blunt such failures. "We could lose 100 nodes in a network of 250,000 and it wouldn't matter," he says.

Scale also brings latency gains. When a network is large enough, every job routes to a nearby device. Far Labs claims latency of 100 milliseconds or less on its platform. Lower cost and latency, the companies argue, could unlock use cases that are currently impractical — such as real-time in-game AI video generation.

"OpenAI last year had US $30 billion in revenue, but they closed the financial year at an $8 billion loss. Why? The official reason is due to the high cost of inference," Shazhaev says. "And those are mostly text models. For gameplay, you have audio, video, animations: It's heavy data, and you need real-time responses. So, we've been trying to solve this issue."

The market stakes are straightforward. If inference cost is what separates AI revenue from AI profit, as OpenAI's reported numbers suggest, a cheaper distributed alternative attacks the industry's core economics rather than its periphery.

For now, these platforms are betting that the world already contains enough idle silicon to matter. "All these big guys are running around building data centers," Shazhaev says, "but I believe there is enough compute power that already exists in the world."

Original: linkedin.com

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

118 articles

Related articles

  1. OpenAI, Oracle and SoftBank Add Five Sites to Stargate Buildout
  2. Modulate raises $25M to detect deepfake voices and read customer intent
  3. A Startup Wants to Train AI World Models on Your Gameplay
  4. Instinct, the iMessage AI Agent, Saves Money and Raises Alarms
  5. Higgsfield Generates 4 Million Videos a Day With GPT-4.1, GPT-5 and Sora 2

Next article »