Introducing Parallel Fast Mode at $1 CPM for Frontier Intelligence

This title was summarized by AI from the post below.

A little about the significance of a $1 CPM, 5-10x more affordable web search product for AIs and agents. 45 days ago we launched a Turbo search mode. Our goal was: how do we get people to think about web search as a no-brainer always-on part of AI and agentic workflows? For many use cases, adding web access can 2-3x eval results! Turbo was unique: very fast (200-300ms) and the first ever $1 CPM AI-focused search SKU. Comparable offerings (including our own) are normally $5-$10 CPM. We felt both would encourage use, but we didn't know which would matter more. Since then, we kept hearing one piece of feedback above all others: "frontier intelligence cost is dropping, and $1 CPM for search makes it a no-brainer for my use case, but I wish I could trade some of Turbo's speed for even higher quality." If frontier intelligence is 5x cheaper than it was a year ago, why isn't search? But there was a reason why deeper search was priced $5+. The search index is large, and growing. Users want fresh data and long tail data. Getting accurate, relevant, and succinct summaries for LLMs to consume requires inference over large parts of the web, often at request time. This costs time and money (compute, inference, index, ...). $1 CPM for frontier quality web grounding felt hard to achieve. But the directive was clear. Parallel Fast mode is the culmination of that effort. It is $1 CPM, just like Turbo, meaning it's a tiny fraction of the total cost of inference, even on cheap models. It's consistently sub-1s (won't add meaningful latency to any LLMs). And most importantly, it's frontier intelligence. No tradeoffs. It's ranked #3 on Artificial Analysis, only 2 points (73 vs 75) below the best search product in the world (our own Advanced mode). Trivially outperforming models at 5-10x the price. For giving AIs and agents web access, we think we created the most obvious default search product. And while we're excited about getting existing search users over to Fast, more importantly, we think it will make instant web access a no-brainer for many more developers.

View organization page for Parallel Web Systems

35,282 followers

Capable models have gotten cheaper. Search hasn’t, until now. Today we’re introducing Fast mode for Parallel Search: frontier-quality web search for AI that’s 10x cheaper than the default from model labs. - $1 per 1,000 results - 700ms p50 latency - #3 on Artificial Analysis Search Index for intelligence Parallel Fast is the only web search that’s cheap, fast, and accurate. It’s optimized for today’s leading class of cost-effective models: GPT-5.6 Luna, DeepSeek V4 Pro, MiniMax M3, and Qwen3.8 27B: - 10x cheaper than frontier labs, 5x cheaper than other APIs at $1 per 1,000 results - On the quality vs. cost Pareto frontier - On the quality vs. latency Pareto frontier - Nearly matches frontier accuracy set by Parallel Advanced (SOTA) Read the announcement: https://lnkd.in/e4PUSb2z

  • No alternative text description for this image

This is huge! Real challenge with buildig with frontier models is, you still need to ground it's responses and the best way to do that is by augmenting it with web search. The built in websearches are sooooo expensive! Love the problem this is solving. Also, they are lucky to have you Vlad Shulman

To view or add a comment, sign in

Explore content categories