Appearance
Model capability, tiers and roadmap
As of 2026-08-12. Written to answer a partner's three questions — what the platform can do, what is coming when, and whether Auto routing is ours to configure — and to plan the video expansion behind it.
Read the honest summary first, because the headline number is easy to misquote.
The one number that matters
| Count | What it means | |
|---|---|---|
| Model rows in the catalogue | 72 | Integrated in code: adapter written, request/response mapped |
| Servable right now | 29 | Model enabled and its provider has a working API key |
| Verified by a real call | 13 | We have actually generated something with it and logged the result |
| Verified failing | 6 | Real call attempted and refused — needs a fix or retirement |
| Unverified | 43 | Never called. Assume nothing about these until they are verified |
Do not tell a partner "72 models". Say 29 live, 13 independently verified, 72 integrated — the gap is not engineering, it is API keys. Only three providers are keyed today (Google/Gemini, OpenAI, ElevenLabs). The other fifteen have finished adapters sitting behind a missing credential, which is why the platform reports them as unservable rather than pretending otherwise.
That is the whole story of this roadmap: most of what a partner is asking for is a procurement task measured in days, not a build measured in months.
1. Feature-wise capability, by tier
Tier is quality-and-cost banding, not vendor prestige:
- Flagship — frontier output, highest unit cost. What you demo.
- Premium — strong, production-grade, mid cost. What most paid usage should land on.
- Standard — fast and cheap. What free and entry tiers run on.
LIVE = servable today · KEY = integrated, waiting on a provider key.
Text generation & chat
| Tier | Model | Status |
|---|---|---|
| Flagship | Claude Opus 4.7 · Gemini 3 Pro | KEY · LIVE (Gemini 3 Pro currently failing verification) |
| Premium | Gemini 3.5 Flash · Gemini Pro (latest) · GPT-4o · Claude Sonnet 4.6 · Kimi K2.6 | LIVE ×3, KEY ×2 |
| Standard | Gemini Flash Lite · GPT-4o mini · Gemma 4 31B · DeepSeek Chat · Grok 2 | LIVE ×3, KEY ×2 |
Vision / multimodal understanding
Same Gemini and GPT-4o line as above — 7 live. Image-in, text-out; used for describe/analyse/OCR-style features.
Image generation
| Tier | Model | Status |
|---|---|---|
| Flagship | GPT Image 2 · Nano Banana Pro · FLUX 1.1 Pro · Ideogram v2 | LIVE ×2, KEY ×2 |
| Premium | GPT Image 1.5 · Gemini 3.1 Flash Image · Gemini 2.5 Flash Image · Stable Diffusion 3 Large · Qwen Image | LIVE ×3, KEY ×2 |
| Standard | GPT Image 1 Mini · FLUX Schnell · FLUX Dev · SDXL · MiniMax Image 01 | LIVE ×1, KEY ×4 |
19 rows, 6 live. Also live: image-to-image, inpaint (mask), and an enhance/upscale path.
Video generation — the thin spot
| Tier | Model | Status |
|---|---|---|
| Flagship | Sora 2 Pro · Veo 3.1 · Runway Gen-4.5 · Kling 2.5 Turbo Pro · Seedance 2.0 | LIVE ×2, KEY ×3 |
| Premium | Sora 2 · Veo 3.1 Fast · Runway Gen-4 Turbo · Hailuo 2.3 · Wan 2.5 · Luma Ray 3.2 · Runway Aleph 2.0 | LIVE ×2, KEY ×5 |
| Standard | Veo 3.1 Lite · Kling 1.6 Pro · Runway Gen-3 · Luma Uni-1 · Hailuo (Runware) · Wan (Runware) | LIVE ×1, KEY ×5 |
27 rows, 5 live (3 Veo + 2 Sora), none verified. Everything else is one API key away. Note Seedance is already integrated twice — as runware/seedance-via-runware and runway/seedance-2.0 — so the partner's ask for Seedance is a signup, not a sprint. (The second row is also mis-filed: Seedance is ByteDance's model, not Runway's. It needs correcting or deleting.)
Beyond text-to-video, already implemented: image-to-video, first-and-last-frame, video extend, and lip-sync (Kling), plus avatar video via Tavus. All KEY.
Speech, music and audio
| Tier | Model | Status |
|---|---|---|
| Flagship | ElevenLabs Multilingual v3 · Eleven Music | LIVE, both currently failing verification |
| Premium | Multilingual v2 · Turbo v2.5 · Gemini 3.1 Flash TTS · MiniMax Speech 2.8 | LIVE ×3, KEY ×1 |
| Standard | Flash v2.5 · Sound Effects · Gemini 2.5 Flash TTS | LIVE ×3 |
| Speech-to-text | ElevenLabs Scribe v1 | LIVE, failing verification |
11 rows, 9 live — the strongest area, and the one with four verification failures to clear.
Also live (platform, not models)
Async job queue with progress and refund-on-failure · prompt enhancement · credit wallet and per-plan model gating · community publish/share/remix · per-model server-side instructions and token budgets.
2. Roadmap
Dated from 2026-08-12. Phases 1–2 are the ones that move the capability numbers, and both are gated on payment details rather than engineering.
Phase 1 — Unlock what is already built (week of Aug 12–19)
| Step | Work | Owner | Effort |
|---|---|---|---|
| 1 | Open self-serve accounts: fal.ai, Runware, Replicate | Business | 1 day, card only |
| 2 | Add credentials in admin (Providers → Add account) | Ops | 1 hour |
| 3 | Correct the upstream model ids on every unverified row | Eng | 1 day |
| 4 | Run the verify sweep, model by model, and read the results | Eng + Ops | 1 day, ~BDT 500 in real calls |
| 5 | Retire or fix the 6 failing rows | Eng | 0.5 day |
Outcome: 29 → ~60 live models, 20 of them video. No new adapter code.
Step 3 is the real engineering: the unverified rows carry placeholder upstream ids (fal/veo-3.1 resolves to veo-3.1, which is not a fal endpoint path). Each needs the vendor's exact id, which is why verification has to follow.
Phase 2 — Video depth and correct pricing (Aug 19–31)
| Step | Work | Effort |
|---|---|---|
| 6 | Price every newly live model from the vendor's real rate card | 1 day |
| 7 | Add a tier field to models so Flagship/Premium/Standard is API data, not a document | 0.5 day |
| 8 | Map tiers onto plans — Flagship on Pro and above, Standard on free | 0.5 day |
| 9 | Storefront: tier badges and per-model sample gallery | Client team |
Step 6 matters commercially: the current catalogue prices are estimates, and video is where a wrong number is expensive — a flagship clip is roughly 20× a standard image in provider cost.
Phase 3 — Direct vendor accounts, if volume justifies (Sep)
Aggregators cost roughly 10–25% more per second than a direct contract and sometimes lag a vendor's newest model by weeks. Once video volume is real, go direct on the two or three models that carry it — Kling (direct also unlocks the lip-sync and extend paths we have already built), Runway, MiniMax/Hailuo. Each is a contract and an invoice relationship: 1–3 weeks of onboarding, no adapter work.
Seedance direct means ByteDance Volcengine ModelArk (CN) or BytePlus (intl), both with business-verification onboarding. Not worth it before volume — take it through fal or Runware first and revisit.
Phase 4 — Auto routing v1 (Sep, 1–2 weeks)
See the next section.
Phase 5 — Avatar, music, STT (Sep–Oct)
Tavus avatar video and MiniMax music are integrated and unkeyed; ElevenLabs music and Scribe are keyed but failing. Together: ~1 week once Phase 1 is done.
What we are deliberately not promising
Timelines above are ours to control. What we cannot promise is a specific third-party model on a specific date — vendors ship, deprecate and re-price on their own schedule, and two of the six verification failures are vendors having changed something under us. The commitment worth making to a partner is the capability: any model reachable over an HTTP API on one of our seven adapter shapes is a catalogue row and a verify run, not a release.
3. Auto routing — can Kikori.ai configure it?
Honest answer: routing exists today at the account layer, not the model layer, and everything that exists is ours to configure.
What is live now
Per request the gateway picks which vendor account serves a named model, and it is entirely admin-driven:
- Weights — split traffic across several keys for one provider.
- Per-account model allowlists — this key serves only these models.
- Daily and monthly spend caps per key, rolled at the period boundary.
- Automatic failover — a rate-limited or exhausted key is marked and the next eligible one takes over mid-request.
- Health-driven exclusion — a provider with no working credential goes INACTIVE and its models stop being offered rather than failing at generate time.
What does not exist yet
An Auto mode that chooses the model for the user. Today the client names a model slug. There is no pseudo-model that reads the prompt and picks Flagship vs Standard.
What we would build (Phase 4, 1–2 weeks)
An auto slug per modality (auto/text, auto/image, auto/video) resolving through an admin-editable policy:
- Eligibility — the caller's plan and the model's live health.
- Intent — prompt length, whether an image is attached, requested duration.
- Policy — cheapest-that-qualifies, best-quality-within-budget, or fastest, chosen per plan by an admin.
- Budget awareness — spend down the wallet on Standard before Flagship.
- Transparency — the response names the model actually used, and the decision is logged so a bad route can be explained.
Fully configurable by us, and by design not a black box: same admin surface as the per-model instructions, no client release to change a routing rule.
4. Where the video models come from
Ranked by how fast they deliver a live model, with the adapter status we already have.
| Source | Adapter | Gets us | Signup | Cost note |
|---|---|---|---|---|
| fal.ai | ✅ done | Seedance, Kling, Veo, Wan, Hailuo and the rest behind one key | Self-serve, card, same day | Per-second, no minimum |
| Runware | ✅ done | Seedance, Kling, Wan, Hailuo — usually the cheapest per second | Self-serve, same day | Prepaid credits |
| Replicate | ✅ done | The long tail, including open-weight video | Self-serve, same day | Per-second; slower cold starts |
| Kling direct | ✅ done (+ lip-sync, extend) | Best Kling rate limits and the features we already built | Contract, 1–3 weeks | Cheaper at volume |
| Runway direct | ✅ done | Gen-4.5, Aleph, Act-Two | Contract, 1–2 weeks | Credit packs |
| Luma / MiniMax direct | ✅ done | Ray 3.2, Hailuo | Contract, 1–2 weeks | — |
| ByteDance (Volcengine / BytePlus) | ❌ not written | Seedance at source | Business verification, 3–6 weeks | Cheapest at real volume |
Recommendation: open fal.ai and Runware this week. Between them the entire flagship video line — Seedance included — goes live without an engineer writing a new adapter, and having both gives failover when one has a bad day. Keep Replicate as the long-tail escape hatch, and revisit direct contracts once we know which two models the traffic actually lands on.
One caveat to carry into every one of these conversations: model names, ids and prices in this document are what our catalogue asserts, and the video market re-prices monthly. Read the vendor's current model page at signup and let the verify sweep — not the catalogue — decide what we tell a customer works.