Architect Launches Liquid Inference, an Auction-Based Router for LLM Inference
Architect Financial Technologies has launched Liquid Inference, a router that auctions every LLM request across competing providers and locks a maximum price before generation begins.
According to MarkTechPost, providers post offers to serve specific models, and each request is auctioned across every provider quoting the named model. The buyer pays the lowest offer that meets its rules. For developers, the change is limited to swapping a base URL while keeping existing code, leaving providers to compete on price.
The request flow has four stages. A client sends a standard OpenAI or Anthropic API call. The buyer's constraints filter which offers are eligible. Providers quoting that model compete, and the lowest qualifying offer wins. The maximum price is locked before generation begins, and billing covers metered usage only.
Buyers can set per-job cost caps, time-to-first-token limits and minimum throughput requirements. They can also require approved regions, zero data retention, and allow lists for providers or models. An Auto mode selects the model for a given unit of work. A LinkedIn post by Harrison adds routing rule presets and full multi-modal support. Account holders can view live order books, per-provider and per-model quotes, and cleared transactions, a level of market data that is unusual for an LLM API.
The router is described as a drop-in replacement for agentic coding tools, with the post listing Claude Code, Codex, OpenCode, Cursor, Pi and Cline. Signup is by email and free. According to Harrison, the first 500 users receive $20 of free inference, and a referral program pays 20 percent of referred fees as free inference plus 10 percent on second-level referrals.
Providers onboard through the Liquid Inference app, and Harrison says new providers are verified "in minutes, not weeks." All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes, and providers can update quotes based on their own costs, which lets them sell spare GPU capacity only when they choose. Payouts run through Stripe with itemized records for every job.
The published comparison sets Liquid Inference against OpenRouter and Hugging Face Inference Providers. OpenRouter routes through price-weighted load balancing using the inverse square of price; Hugging Face defaults to the fastest provider with an optional :cheapest suffix. Liquid Inference lists "hundreds" of models but does not disclose its providers, while OpenRouter cites more than 500 models across 80-plus providers and Hugging Face lists 18 partners. Liquid Inference is compatible with both OpenAI and Anthropic APIs; OpenRouter is OpenAI-compatible, and Hugging Face is OpenAI-compatible for chat only. Liquid Inference locks a maximum price before the first token, OpenRouter exposes a max_price parameter, and Hugging Face does not disclose a price cap. OpenRouter charges a 5.5 percent card credit fee with a $0.80 minimum, Hugging Face applies no markup, and Liquid Inference has not disclosed its platform fee. Liquid Inference publishes live order books and cleared trades, which the other two do not disclose.
Architect's background is in trading infrastructure. In May 2026 the firm acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review, and it used that exchange-building experience to construct the inference marketplace. Fees, the provider list and latency data are not yet public.