How do I control what data AI can see?
Define exactly what data and systems AI and agents can access—and what stays private.
One shared URL. Everyone sees the same task, progress, decisions, and context. Teammates can correct the agent, approve a step, or hand the Session to someone else while the work is still running.
Define exactly what data and systems AI and agents can access—and what stays private.
Qua’s Answer Waterfall finds the lowest-cost path for every question.
Qua searches your private knowledge first. Most answers come from what you already own.
Agent Sessions turn AI into a shared workspace—see it, edit it, approve it, together.
Use any model, in any infrastructure. One policy layer stays in control.
The most expensive tier sits at the top. Qua resolves every question as far down the waterfall as it can — 84% never reach the most expensive tier, and Open Models are only ever used when you ask for them.
The highest-capability models, for complex reasoning and premium tasks — reached only when the expected-value gate clears.
Pro Models · premiumOptimized for speed and cost — the cheapest model with a proven acceptance rate for the task class.
Fast Models · meteredOpen-weight and open-source models, available whenever you ask for one. Qua never routes here on its own.
By request only · never auto-selectedSearches connected company data and approved workplace sources — including org-verified answers — then a cited answer.
≈95%+ cheaper than external models*
In-perimeter · ~$0Retrieves what you saved or connected — your verified answers, scoped knowledge, and instant deterministic resolution.
$0 · instantBelow is a fully modeled 30-day scenario for a 380-seat enterprise — our Meridian Group demo dataset. The same 286,400 queries are priced three ways: what they cost through Qua's Answer Waterfall, what an all-open-source setup would cost, and what routing everything to a premium model would cost. Every figure on this page reconciles — totals, tiers, and daily series all add up.
84% of work never reaches the most expensive tier.
Frontier models process 6% of the tokens but half the bill — routing is the entire game.
Illustrative modeled scenario (Meridian Group demo dataset), not customer results. The open-source and all-premium figures are modeled counterfactuals at list rates; only the Qua figure represents an actual invoice in the scenario. Labor-value figures, where shown, use a blended $85/hr rate.
*Measured routing outcome, varies by workload. Independent research on model routing reports 40–98% cost reduction: RouteLLM (UC Berkeley, ICLR 2025) measured 85% at 95% of GPT-4 quality; FrugalGPT (Stanford) up to 98%.
Qua's documented 5× total-spend reduction is equivalent to roughly 80% lower spend versus an ungoverned baseline. It should be presented as a measured routing outcome, not as a universal RAG claim.
Major model platforms document input-token savings of up to 90% when repeated context qualifies for prompt caching. Retrieval and output generation still carry cost.
Qua is model-agnostic and hosting-agnostic: closed commercial models, open-weight models, or your own fine-tunes — running in Qua cloud, inside your VPC, or fully on-prem.
Zero-cost-tier answers — Personal Knowledge and Enterprise Search — are composed from your own knowledge and never leave your perimeter.
Neutrality is the business model: Qua never trains shared models on your knowledge, and admins see outcomes — never prompt content.