On this page
AI Search Prompt Clusters: Architecture and True Costs
tl;dr
RAILS achieved 59.3% clustering accuracy across six public benchmarks by replacing traditional embedding pipelines with single LLM prompts, beating the strongest prior LLM clustering method. The system replaced Zendesk's production HDBSCAN stage, and practitioners must account for hidden infrastructure and LLM costs beyond headline API pricing.
A retrieval-augmented clustering system called RAILS just lifted clustering accuracy from 51.2% to 59.3% across six public benchmarks — and it did so by throwing out the embedding pipeline entirely and packing candidate labels into a single LLM prompt. If you work anywhere near AI search prompt clusters — the grouping of queries, tickets, or keywords into themes that downstream systems act on — that result should reframe how you think about the whole stack. The interesting story isn’t just the accuracy jump. It’s what the architecture, the infrastructure, and the billing actually look like once you peel back the headline numbers.
What are AI search prompt clusters, exactly?
AI search prompt clusters are thematic groupings of queries or prompts — support tickets rolling up into topic clusters, search queries fanning out into sub-queries, keywords consolidating into content hubs. The clustering layer decides what your downstream products see, so its quality bounds everything built on top of it.
This layer is quietly becoming regulated, patented infrastructure. On August 31, 2026, the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act — the first time an AI chatbot got that classification, triggered by the fact that it retrieves from the live web and shapes what users see. When Brussels treats your chatbot as a search engine, the mechanics of how queries get grouped and routed stop being an implementation detail.
Meanwhile, the vendors are fencing off the techniques. Perplexity’s augmented search engine patent, granted in September 2026, describes generating a user prompt from an initial query, storing it in a search state database, and using that state to generate follow-up search queries and summaries. That’s a prompt-clustering loop, patented. The takeaway: this isn’t a niche ML problem anymore. It’s core plumbing with legal and strategic weight.
How does RAILS cluster tickets at production scale?
The short answer: retrieval over a growing label pool, not embeddings over the whole corpus. RAILS is a retrieval-augmented incremental LLM clusterer that retrieves top-K candidate labels from a per-brand pool and packs them into a single LLM prompt alongside a batch of tickets. The model returns one of three outcomes per ticket — reuse an existing label, create a new one, or mark it as noise — and new labels get absorbed back into the pool through threshold-based merging.
The benchmark results are the strongest corroborated data point in this space right now. Per the RAILS paper from Zendesk’s team, the system hit 59.3% accuracy, 74.8% NMI, and 54.7% ARI on six public benchmarks, beating the strongest prior LLM-clustering method on average. Those aren’t vendor marketing numbers; they’re from an arXiv paper with named authors and a described methodology.
Here’s the part that matters for practitioners: RAILS replaced a traditional HDBSCAN stage in Zendesk’s production ticket-topic-discovery pipeline, after the team ran an embed-then-cluster pipeline across hundreds of millions of tickets for over a year. That’s a real production migration, not a lab demo. The embed-then-cluster approach you’d reach for by default — encode, reduce with UMAP, cluster with HDBSCAN — accumulated enough structural limitations at scale that its own operators replaced it. If you’re still defaulting to that pattern, it’s worth asking why.
Why do zero-shot query fan-outs collapse?
Because generic LLMs aren’t search experts, and no amount of prompt engineering fully fixes that. Google Research’s Retrieve-for-Train framework uses offline reinforcement learning to train a lightweight diffusion model that generates reward-aligned query fan-outs at inference time — bypassing the expensive test-time reasoning budget entirely.
The problem it solves has a name worth knowing: paraphrastic collapse, where zero-shot LLMs asked to decompose a broad query generate redundant near-synonymous sub-queries instead of complementary facets. Ask a stock model to fan out “camping gear” and you get ten variations of “four-person tent” rather than a coherent slate covering tent, sleeping bag, stove, and headlamp. The model isn’t navigating your corpus; it’s just predicting text.
This connects to a pattern we’ve tracked in AI prompt patterns for developers: zero-shot behavior is task-dependent, and the tasks where it fails hardest are the ones requiring set-level properties like diversity and coverage. You can’t prompt your way to corpus-aware decomposition. You either train for it (Retrieve-for-Train’s approach) or retrieve your way around it (RAILS’ approach). Both paths beat throwing a bigger thinking budget at a general-purpose model.
What does an AI search prompt cluster stack actually cost?
Here’s where a pattern I’ve observed kicks in — call it cost unbundling. The headline price in any AI pricing table is systematically lower than your total cost of ownership, because the real spend lives in adjacent layers: the LLM bill behind the seat license, the RAM behind the vector store, the engineering time behind the “free” tier. Let’s price the actual layers.
| Layer | Tool | Pricing | Billing model |
|---|---|---|---|
| Search API | Perplexity Search API | $5 per 1,000 requests | Per-request |
| LLM tokens | Perplexity Sonar | $1 to $15 per million tokens | Token-metered |
| LLM tokens | DeepSeek V4 Pro | $0.66/M input off-peak, $1.32/M at peak; $1.98/M output off-peak, $3.96/M at peak | Token-metered, time-of-day |
| LLM tokens | Claude Opus 5.5 | $4/M input, $20/M output, cache reads at $0.20/M | Token-metered |
| Vector storage | Qdrant Cloud | RAM-based; ~3 GB RAM per million 768-dim vectors | Resource-provisioned |
| LLM proxy/governance | LiteLLM Enterprise | Commonly cited at ~$250/month Basic, ~$30,000/year Premium | Self-hosted license |
A few things stand out. First, the cheapest per-token option isn’t automatically the cheapest stack. DeepSeek V4 Pro’s MIT-licensed weights are a portability guarantee, not an immediate saving — self-hosting a 1.7T-parameter mixture-of-experts model requires GPU capacity most teams don’t have sitting idle. Second, Qdrant’s RAM-based pricing means your cost driver is vector footprint, not query volume: a 768-dimensional float32 vector occupies roughly 3 KB, so capacity planning looks like traditional database provisioning. The Qdrant Foundations plan is genuinely free — unlimited clusters up to 250 cores combined, unlimited users, 15-day metric retention — which makes it the right starting point for most cluster pipelines.
Seat pricing has its own traps. Perplexity Teams and Enterprise start at $40 per month or $400 per year per seat, and the free tier’s 3 Pro searches per day and 1 Research query per month won’t survive a week of real research work. At 50 seats on annual billing, Perplexity Enterprise Pro runs $20,000 per year against Claude Team’s $12,000 — but that comparison only holds if the job is cited research rather than analysis work on documents you already have. And LiteLLM’s tiers are self-hosted, so the license fee excludes the PostgreSQL, Redis, patching, and on-call rotation you now own. The sticker is never the bill.
How do content teams cluster prompts differently?
They cluster by what Google actually ranks, not by what text looks similar. The SEO Cluster tool groups keywords by Google SERP overlap — shared top-10 results — rather than text similarity, then designs hub-and-spoke content clusters with internal link matrices. That’s a fundamentally different clustering signal: two keywords with zero lexical overlap belong in the same cluster if Google ranks the same pages for both.
This matters more than it used to, because AI search visibility varies drastically by query category — as we found in our analysis of AI search visibility by query category, citation mechanics differ sharply across query types, so a blended visibility score is close to meaningless. If your clustering layer lumps category queries, how-to queries, and evaluation queries into one bucket, your content strategy inherits that confusion.
The practical implication: match your clustering signal to your distribution channel. SERP overlap for search-driven content, retrieval-augmented LLM labeling for support tickets, trained fan-out models for database-aware search. Text similarity alone is the weakest signal of the three, and it’s the one most teams default to because it’s the easiest to compute.
Which clustering approach should you pick?
Match the architecture to your constraint, not to the benchmark leaderboard:
- High-volume, evolving label spaces (support tickets, feedback streams): the RAILS pattern — retrieval over a per-brand label pool with incremental merging. It’s proven in production at Zendesk scale, and the prompt-driven control means you can steer clusters without retraining anything.
- Search fan-out over a fixed corpus: don’t zero-shot it. Paraphrastic collapse will eat your coverage. Either adopt a trained approach like Retrieve-for-Train or constrain the decomposition with retrieval signals from your actual index.
- Content architecture: SERP-overlap clustering, full stop. Text similarity tells you how words relate; SERP overlap tells you how Google thinks they relate. Only one of those drives rankings.
- Budget-constrained experimentation: start on Qdrant’s free Foundations tier and DeepSeek’s off-peak API pricing, but model the RAM growth and peak-hour exposure before you commit. The free layers are real; they’re just not where the cost story ends.
One open question worth tracking: as clustering moves from embeddings to LLM prompts, measurement hasn’t caught up. Only a minority of brands track AI search performance systematically — our piece on how to measure AI search visibility argues infrastructure-embedded measurement beats synthetic prompt tools — and I’d bet the same gap exists inside clustering pipelines. If you can’t measure cluster quality against downstream outcomes, you’re optimizing the benchmark, not the business. Start there.
Recommended Reading
-
Prompt CI/CD Explained: Why Traditional Pipelines Fail LLMs
You should decouple prompt updates from code deploys because traditional CI fails LLM applications: prompts change faster than binaries and fail silently. 70% of teams update prompts monthly and 10% daily, yet CircleCI finished dead last at 13 minutes 18 seconds versus Semaphore's 5 minutes 1 second, proving speed branding is decoupled from runtime reality.
-
How AI Search Finds and Recommends SaaS Products
44% of B2B SaaS products are functionally invisible to AI buyers, with most purchase decisions now made via AI-generated shortlists before any sales contact. This post breaks down the 'proof density' ranking signal AI search uses, why legacy ABM tools fall short, and how to optimize for AI-driven discovery to capture pipeline.
-
AI Search Consensus Signals: What Actually Drives Visibility
Only about 12% of AI-cited URLs rank in Google's top 10, making a dedicated AI visibility strategy essential. Success depends on consensus signals across independent sources, since most brand mentions in AI answers come from third-party pages.