Products & Tools

Trading Firm Architect Launches Liquid Inference, an Auction Market for LLM Requests

Architect Financial Technologies launched Liquid Inference, auctioning every LLM request across competing providers so buyers pay the lowest qualifying offer, with max price locked before the first token.

Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference
Architect Launches Liquid Inference, a Real-Time Auction for LLM InferenceAI-generated
By Marcus Bennett5 min read

Updated

Why it matters

  • Architect Financial Technologies launched Liquid Inference, an exchange-style router that auctions every LLM request across competing providers.
  • First 500 users get $20 of free inference; referrals earn 20% of referred fees as free credits plus 10% on second-level referrals.
  • Buyers can set per-job cost caps, time-to-first-token limits, minimum throughput, approved regions, zero data retention, and allow lists.
  • Architect acquired a US Designated Contract Market in May 2026 to list GPU compute futures, pending regulatory review.
  • OpenRouter lists 500+ models and 80+ providers versus Liquid Inference's 'hundreds' of models with an undisclosed provider count.

Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every single request, with providers bidding against each other to serve each prompt and the buyer paying the lowest offer that meets its rules. The pitch to developers is blunt: swap a base URL, keep your code, and let providers compete on price.

The launch matters because it imports exchange-style price discovery — the core mechanism of commodities and derivatives markets — into the market for AI compute, where pricing today is largely a posted-rate affair set unilaterally by model providers. Architect is not an AI lab; it is a trading firm that runs the AX perpetual futures exchange and acquired a US Designated Contract Market in May 2026 to list GPU compute futures, pending regulatory review.

What is Liquid Inference?

Liquid Inference is an exchange-style router for LLM inference. According to Architect, providers post offers to serve specific models. Each request is auctioned across every provider quoting the named model. The lowest-priced offer that meets the buyer's rules wins the job.

The product draws directly on the team's background building financial exchanges. Architect applied that experience to create two-sided price discovery for inference: buyers with constraints on one side, competing sellers of GPU capacity on the other, and a clearing mechanism in between.

That lineage separates it from typical AI routing services. Rather than a static rate card or a load balancer, the platform treats each prompt as a tradeable unit of work with a market-clearing price.

How does the inference auction work?

The flow has four steps:

  1. Request: A client sends a standard OpenAI or Anthropic API call.
  2. Rules: The buyer's constraints filter eligible offers.
  3. Auction: Providers quoting that model compete, and the lowest qualifying offer wins.
  4. Receipt: The max price is locked before generation, and billing covers metered usage only.

Buyers can set per-job cost caps, time-to-first-token limits, and minimum throughput, according to Architect's announcement. They can also require approved regions, zero data retention, and provider or model allow lists. An Auto mode can pick the model for a given unit of work. A LinkedIn post by Harrison adds routing rule presets and full multi-modal support.

Account holders can view live order books, per-provider and per-model quotes, and cleared transactions. That level of market data is unusual for an LLM API, where users normally see nothing beyond a monthly invoice.

What do buyers get?

The platform targets developers already using agentic coding tools. The Harrison post lists drop-in compatibility with Claude Code, Codex, OpenCode, Cursor, Pi, and Cline. Onboarding is a free email signup, and the first 500 users get $20 of free inference, per Harrison.

A referral program adds 20% of referred fees back as free inference, plus 10% on second-level referrals.

What do inference providers get?

Providers onboard through the Liquid Inference app. Harrison says new providers are verified "in minutes, not weeks." All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes.

Providers can update quotes based on their own costs, which lets them sell spare GPU capacity only when they want to. Payouts run through Stripe, with itemized records of every job.

How does Liquid Inference compare with OpenRouter and Hugging Face?

The closest existing services route differently. OpenRouter uses price-weighted load balancing — weighting by the inverse square of price — across 500+ models and 80+ providers. Hugging Face's Inference Providers default to the fastest provider among its 18 listed partners, with a :cheapest suffix as an option. Liquid Inference instead auctions every request across all providers quoting a model, claiming "hundreds" of models; its provider count is not disclosed.

Feature Liquid Inference OpenRouter Hugging Face Inference Providers
Routing model Per-request auction across quoting providers Price-weighted load balancing, inverse square of price Fastest provider by default; :cheapest suffix optional
Model / provider count "Hundreds" of models; providers not disclosed 500+ models, 80+ providers 18 listed partners
API compatibility OpenAI and Anthropic-compatible OpenAI-compatible OpenAI-compatible, chat only
Price cap Max price locked before first token max_price parameter Not disclosed
Data controls ZDR, regions, allow lists zdr, data_collection, only/ignore Provider preference order
Platform fee Not disclosed 5.5% card credit fee, $0.80 minimum No markup
Free credits $20 for first 500 users Not disclosed $0.10/month free, $2.00 PRO
Public market data Live order books and cleared trades Not disclosed Not disclosed

The Anthropic API compatibility is a differentiator: OpenRouter and Hugging Face are OpenAI-compatible only, while Liquid Inference supports both OpenAI and Anthropic clients.

What are the open questions?

Liquid Inference's platform fee is not disclosed. The provider list is not disclosed. Latency data is not public. Until those details arrive, buyers cannot fully price the tradeoff between auction-driven savings and the overhead of an intermediary auction on every request.

The exchange framing also raises a market-structure question. Architect is building a compute futures business — its May 2026 Designated Contract Market acquisition, pending regulatory review, targets GPU compute futures — and a spot market for inference would give that business a natural price reference. Whether the auction produces liquid, trustworthy prices at scale will determine whether inference becomes a genuine commodity market or a router with marketing.

The first 500 users' $20 of free inference is the cheapest way for developers to find out.

Original: liquidinference.ai

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

150 articles

Related articles

  1. Modal Labs Nears $750M Round at $15.75 Billion Valuation
  2. Amazon open-sources Strands Decider 2B as decision models multiply
  3. Mirakl Bets on AI Agents and ChatGPT Enterprise to Rebuild Commerce
  4. OpenAI and Broadcom Unveil Jalapeño, a Custom LLM Inference Chip
  5. OpenAI and PwC Team Up to Rebuild the Office of the CFO Around AI Agents

« Previous article