Build your own video platform
The long-form guide to shipping an agentic or editor-style video product on the API: the document format, the command reducer, previews, renders, costs and both auth models.
A guide for a developer who wants to ship a video product without building a renderer.
Everything below was checked against the live API before it was written. Field names, credit costs, scope names and status strings are what the server actually returns, not what a spec says it should. Where something is not built yet, it says so.
Base URL: https://api.vidmoat.com/v1. (/api/v1 resolves to the same place, but use /v1: it is the supported path, and it avoids a redirect, which is where Authorization headers get dropped.)
1. What Vidmoat gives you, and what it does not#
You get#
- A document format. A project is one JSON document: settings, tracks, clips. Every clip field that affects the output is in it: transforms, crops, keyframes, colour grade, effects, masks, audio DSP, transitions. It is readable, diffable, and you can hold it in your own database if you want to.
- A command reducer. ~80 named operations (
addClip,sliceClip,addCaptions,cutRanges,autoDuck,addKeyframe,applyMotionPreset,addAdjustmentLayer, …) that take a document and produce a new one. Published verbatim atGET /v1/schema/commands. - A render farm.
POST /v1/rendersqueues a job; a worker composites the document in Chromium, mixes audio through ffmpeg, and hands you an MP4/WebM/GIF URL. You never install ffmpeg. - Generative media. Text-to-image, text-to-video, text-to-speech, AI stickers, behind a credit meter, with refunds when the provider fails. xAI Grok is the default; images can also route to OpenAI or Vertex, speech to OpenAI, and video to HappyHorse or Wan.
- Transcription. Word-level timings in exactly the shape the
addCaptionscommand consumes. Plus a free mode that distributes a script you already have across a known duration, with no provider call at all. - Preview. Three representations: cheap JSON metadata + layout lint, a single rendered JPEG frame, and the whole interactive composition as HTML you can scrub client-side at zero server cost.
- A planner.
POST /v1/ai/agentturns plain English into a validated command program and applies it.
You bring#
The product. The UI. The users. The onboarding, the pricing page, the retention loop, the reason anyone opens your app instead of CapCut.
What is not there#
Say these out loud before you design around them:
- Render webhooks are available in the developer console. Subscribe to
render.completedandrender.failed(plus credit, app and Telegram bot events), verify signatures and inspect delivery logs at/developer/webhooks. Public/v1/webhooksmanagement is not shipped. Poll/v1/ai/jobs/{id}for video generation jobs and keep bounded render polling as a recovery path. See the webhook guide. ai.analyzeis not exposed. The scope exists. The analysis (shot boundaries, loudness, voice activity) runs internally. There is noPOST /v1/ai/analyzeroute. Do not build a cut-decision engine that assumes you can ask Vidmoat what is in a frame.- No idempotency keys. There is no
Idempotency-Keyheader. Commands are not idempotent: postingaddCliptwice gives you two clips. §8 explains what to do instead. - No chunked upload.
POST /v1/media/uploadsbuffers the whole file in memory and refuses anything over 32 MB. Bigger files go throughPOST /v1/media/import, which streams from a URL you host (up to your plan's upload cap or 512 MB, whichever is smaller, with a 120-second timeout). - No browser-facing API. An API key is a server credential. It must never reach a browser, and CORS on
api.vidmoat.comis configured for Vidmoat's own hosts, not for arbitrary third-party origins. Your web UI talks to your server; your server talks to Vidmoat. §4 shows the one place this matters architecturally. - No frame-exact local scrubbing. The interactive preview is a real browser playing your media over the network, driven by the page's own clock. It is excellent for composition and timing decisions. It is not a broadcast monitor. If your users are colour-grading 4K ProRes and need guaranteed frame accuracy with no network, Vidmoat is the wrong choice: you want a local NLE, and no amount of API is going to fix that.
- No social publishing. No scope grants posting to Vidmoat Social, and none will.
- Browser recorder, inpainting and Flows are deliberately not in v1. Each would let one API caller starve everyone else on the box.
2. The mental model#
This is the one section worth reading twice, because everything else follows from it.
A project is a JSON document.
Nothing is derived and stored. The timeline's duration is max(clip.start + clip.duration), computed on read. Compositing order is the position of a clip's track in the tracks array: a clip whose track is listed later draws on top. That is a real gotcha: trackIndex is a track id, not a layer number, and the two only coincide in the default document.
Every edit is a command through one reducer.
One reducer is the only thing in the system that mutates a document. It generates ids, allowlists patch fields, normalises clips, resolves collisions, and applies the auto-fixes that model output actually needs.
The same reducer serves the editor, the agent and the API. The editor's window.vidmoat.dispatch calls it. The chat agent calls it. The MCP server's edit_project calls it. POST /v1/projects/{id}/commands calls it directly, as a library import, not over HTTP. PATCH /v1/projects/{id} calls it, even for a rename, because Project.name on the database row and document.settings.name are two copies of one fact and setProjectSettings is the only thing that keeps them equal.
Why this matters to you. An API edit and a hand edit cannot diverge. If your user opens the same project in Vidmoat's own editor, they see exactly what your API calls produced, with the same ids, the same normalisation and the same semantics. There is no "API version of the document" and no import/export step where fidelity leaks. That is the guarantee you are buying, and it is the reason a third party can build a real editor on this at all: you are not driving a sanitised subset, you are driving the actual thing.
Two consequences to design around:
- Read the document you are about to mutate.
GET /v1/projects/{id}?view=documentreturns it verbatim: the same JSON the reducer consumes. Nothing is reshaped, so your read and your write agree about field names. - The reducer may not do exactly what you asked. An overlapping clip is pushed to the next track index, and the track is created if it is missing.
addClipdefaultsmediaDurationtodurationif you omit it, which quietly breaks waveform windowing and looping. Always read the returnedresults[].data.clipIdrather than assuming, and passmediaDurationexplicitly when you passmediaStart.
3. Hello, render#
3.1 Get a key#
At https://developer.vidmoat.com/apps: create an app, accept the developer agreement, mint a key with the scopes you need.
- Key format:
vmk_live_<48 hex>orvmk_test_<48 hex>. Older unprefixedvmk_<hex>keys still work. - Live keys need three things: the developer agreement accepted, the app approved in review, and a plan with
api_access(Studio, Team or Enterprise). Test keys are free on every plan, including free. Whether a key is test or live is fixed by how it was minted. Max 10 active keys per app. - Scopes are chosen at mint time and clamped to the app's approved ceiling, and clamped again on every request, so narrowing an app's ceiling immediately narrows every key it already issued.
- Keys cannot mint keys.
/api/developer/**is session-authenticated only.
Everything is Authorization: Bearer <key>.
3.2 Who am I#
credential.scopes is the answer to every 403 you are about to get. Read it first.
3.3 Create a project, already edited#
POST /v1/projects accepts a commands array, so a whole video can be one call. Returns 201.
Response:
If a command in the program fails, the project is not rolled back. Each command succeeds or fails on its own, and the ones that succeeded are kept. You get the project plus a results array telling you which command was wrong: a debuggable state, rather than burning a project-quota slot on every retry.
3.4 Apply more commands#
Notes that will save you an afternoon:
- 1 to 200 commands per request. Split larger programs; they apply in order.
- There is no all-or-nothing rollback. Each command succeeds or fails on its own in
results[](ok: falseplus anerror), and the batch is saved if at least one command succeeded. The top-levelokis true only when every command succeeded. blockedis not an error. Plan-gated ops (Pro effects/filters/ transitions,runScript) are dropped from the batch and named inblocked; everything else applies. A Creator-plan caller with one Studio-only effect in a fifty-command program gets forty-nine applied.lintcomes back on every call and you should read it. Overlapping text, off-canvas elements, font sizes unreadable at output resolution.severityis"warn"or"note", capped at 10 entries. Numeric x/y/fontSize choices routinely look wrong on the real canvas, and this is the only cheap way to find out.dryRun: trueruns the whole reducer and returns results + lint without persisting: nothing is saved.
3.5 Render and poll#
202, and note what is not in that request body: the document. You name a project you own and the server reads it. The internal /api/render takes the whole timeline inline; v1 deliberately does not, because that would weld the clip format into a public contract.
applied tells you the two things that would otherwise be silent surprises: a plan-capped resolution (maxHeight + cappedByPlan: true) and a baked-in watermark (free plans lack no_watermark).
Status strings are uppercase: PENDING, PROCESSING, COMPLETED, FAILED, CANCELLED. The failure reason is on render.error (not errorMsg, which is the database column name; the serializer renames it).
Request-shape validation happens for accepted bodies: format ∈ {mp4, webm, gif}, quality ∈ {draft, standard, high}, resolution is "full" or anything else (which means half). A project with no clips is a 400.
4. Building an editor UI#
This is the section that decides whether you have a product or a script. The answer is yes, you can build a real editor UI on this, and here is exactly how.
4.0 The architecture constraint, first#
Your API key is a server credential. It cannot go in a browser, and api.vidmoat.com's CORS configuration is not written for arbitrary third-party origins. So:
The control plane goes through your server. The media plane does not: /uploads/, /exports/ and /fixtures/ on api.vidmoat.com are all served with Access-Control-Allow-Origin: *. That single header is what makes a genuine third-party editor possible, because it means your browser can fetch() the raw media, decode it, and read pixels and samples out of it without tainting a canvas or proxying gigabytes through your own box.
Give yourself a thin proxy and never think about it again:
4.1 Read the document, draw a timeline#
Three views exist and the difference is size, not taste:
view | Returns | Use for |
|---|---|---|
summary | id, name, timestamps, durationSec, clipCount, settings | project lists |
clips (default on a single GET) | + tracks, + per-clip summaries (id, type, start, duration, trackIndex, src, first 120 chars of text, effect types, keyframedProps) | a timeline overview |
document | + the whole document verbatim | anything you intend to edit |
view=document is refused on the list endpoint (GET /v1/projects): a page of 20 full documents is megabytes. You get view=clips and a note saying so.
Drawing the timeline is then arithmetic. Pixels-per-second is your zoom; the Vidmoat editor uses 10 to 300 px/s with a 46 px track height:
4.2 Waveforms, client-side#
This is the part people assume needs a server. It does not, and Vidmoat's own editor does not use one either: it decodes the media in the browser with AudioContext.decodeAudioData and reduces it to peaks. Because /uploads/ sends Access-Control-Allow-Origin: *, your origin can do the identical thing.
A trimmed clip must show its region of the file, not the whole file's. The window is a pair of fractions derived from mediaStart, duration and speed:
The same technique gives you the other analyses Vidmoat runs client-side, off the same decoded buffer: silence detection (RMS per 20 ms window, merge, drop regions under 0.4 s) and energy-flux beat detection. Both are a few dozen lines over the same peaks array. The point here is that none of it costs you a server round-trip or a credit, and it is the input to the cutRanges, trimSilence, addMarkers and sliceClip commands.
4.3 The composition preview#
GET /v1/projects/{id}/preview?format=html returns the whole composition: a self-contained HTML document with every clip as a positioned element, driven by one paused GSAP timeline registered at window.__timelines.main. This is the exact HTML the render worker drives, not an approximation of it.
Two things stop you from putting that URL straight in an <iframe src>:
- You cannot attach an
Authorizationheader to an iframe navigation. - Responses from the API host carry
X-Frame-Options: DENYandContent-Security-Policy: frame-ancestors 'none'.
So fetch it on your server and serve it from your own origin. The HTML uses root-relative URLs (src="/uploads/…", <script src="/vendor/gsap.min.js">), which would resolve against your origin, so rebase them:
Then scrub it:
Wire that to your playhead and you have an interactive preview at zero server cost. Structure you can rely on inside the document:
#rootcarriesdata-composition-id="main",data-start,data-duration,data-width,data-height,data-fps.- Every timed element has
class="clip"plusdata-start,data-duration,data-track-index. Visibility is driven by the timeline;.clipstartsvisibility: hidden. - Video clips are
<video id="media-{clipId}" muted playsinline>. A video's sound is a separate<audio id="aud-{clipId}">element, so muting picture and sound are independent.
One real limitation, stated plainly. The timeline tweens a video's currentTime only when the clip is non-trivially timed: speed !== 1, reverse, freezeAt, or loop. At speed 1 it does not, because the render worker seeks each video explicitly instead. So a bare seek(t) shows the right layout at t but may show the wrong frame of a normal-speed video. Do what the worker does:
data-media-start on the element carries the same number if you would rather read it out of the DOM than out of the document.
4.4 Frame previews, for thumbnails#
The image response carries X-Vidmoat-Frame-Time, X-Vidmoat-Frame-Width, X-Vidmoat-Frame-Height, so you do not need a second JSON call. The JSON response gives you frame.url (durable, for a UI) and frame.dataUri (bytes, for feeding a vision model without a second authenticated fetch).
Budget carefully:
- Every frame launches Chromium, serialised process-wide, and is rate limited to 10 per minute per key in its own bucket (
api_v1_frame), separate from your plan RPM. Exceeding it is a429whose message tells you to useformat=htmlinstead. - Rendered at half the project resolution. Enough to judge composition.
atpast the end is clamped, not rejected.- Omitting both
atandformat=imagegives you the cheap answer: duration, canvas,clipCount,lint,suggestedPreviewTimes, and the composition URL. No browser is launched. This is the call to make on every edit.
Do not build a filmstrip out of this endpoint. Ten frames a minute is a verification tool, not a thumbnail service. For a filmstrip, draw the video element to a canvas yourself in the browser. /uploads/ sends ACAO: *, so the canvas is not tainted and toDataURL works.
4.5 Optimistic edits, reconciled#
The reducer is deterministic, so you can apply an edit locally and post the same command. What you must not do is assume your local result is the server's: collision handling may have moved the clip to another track, plan gating may have dropped an op, and ids are generated server-side.
For anything a user can spam (dragging a clip, scrubbing a slider), debounce into one command rather than one per frame: you have 20/120/600 requests per minute depending on the owner's plan (Team and Enterprise get 600), and 200 commands per request is plenty of room to batch.
5. Building an agentic platform#
Two routes. They are not competitors; they answer different questions.
5.1 Route A: Vidmoat's planner does the thinking#
Body: { projectId, message, dryRun? }. message is required (max 2000 chars), projectId is required. Scope ai.agent, plan feature ai_console (every plan, including free), 5 credits.
What it actually is, so you can price and reason about it:
- One turn, not a loop. One planning call plus at most one repair round-trip, then apply and return. The editor's agent runs a multi-step goal; this does not. Drive the loop yourself by reading
results/unresolvedand calling again. An endpoint that silently makes five provider calls is one you cannot budget. - Provider chain: DeepSeek → Grok → Claude, whichever are configured, plus a built-in offline heuristic planner if none are.
plannerin the response tells you which answered ("heuristic"means no model provider is configured on the deployment). - The repair round-trip is only accepted if it is better: strictly fewer validation failures than the first attempt. Anything it could not fix comes back in
unresolved, rather than being swallowed. dryRun: truereturns the plan without applying it. This is the interesting mode: see what the model intends, filter it, and post the parts you want to/commandsyourself. The 5 credits are still spent: the model call happened, and that is what you paid for.- No privileged path for model output. The commands go through the same reducer and the same plan gate as anything else. A hallucinated Studio-only effect on a Creator plan is dropped into
blocked. - Empty plan is not refunded. A model call happened.
commands: []with a message saying it could not turn that into commands. The 5 credits come back only when every planner fails outright, which is a502 provider_error.
Use Route A when the natural-language step is your product's interface and you do not want to own a model relationship.
5.2 Route B: your own LLM emits commands#
No scope required (a valid credential still is). This is COMMAND_SCHEMA verbatim: the same constant the editor, the chat agent, Flows and the MCP server read. There is no second copy to drift.
The dependable pattern:
Why validate-then-apply matters, concretely. POST /v1/projects/{id}/validate needs only projects.read. That is the whole reason it is a separate endpoint rather than a flag on /commands: a read-only key (the default the portal offers) can check a generated program against a real document without any write access. A self-correcting loop that mutates the project to discover whether its plan parsed is a loop that leaves debris behind on every failed attempt.
It also separates two answers a single boolean conflates:
results[]: is each command valid against this document? Real clip ids, in-range times.blocked[]: does the owner's plan allow it? A Creator-plan caller whose Studio-only effect is perfectly well-formed should be told exactly that, not handed a validation failure to debug.
And lint[] is the third answer, the one nobody asks for and everybody needs: a program can be entirely valid and entirely allowed and still put your title off-canvas. Feed lint warnings back to the model the same way you feed validation errors. suggestedPreviewTimes then tells you which timestamps to actually look at.
5.3 Which route#
Route A /v1/ai/agent | Route B your LLM + /validate | |
|---|---|---|
| Who owns the prompt | Vidmoat | you |
| Cost | 5 credits per turn, predictable | your model bill |
| Latency | one or two provider calls | yours |
| Steerability | message only | total |
| Model choice | Vidmoat's chain | yours |
| Good for | "describe your edit" as a feature | a product whose editing intelligence is the differentiator |
Nothing stops you using both: Route A for a quick-action bar, Route B for your core pipeline.
6. Auth: two models#
Model 1: API key#
Your server acts as itself. One Vidmoat account (yours) owns every project, every byte of media and every credit spent. Your users never hear the word Vidmoat.
- You own the user relationship completely. No consent screen, no revocation you do not control, no OAuth to debug.
- You pay. Every render comes out of your export allowance, every generation out of your credit balance. Your unit economics are your problem, which is also the point: you can charge whatever you like on top.
- Every project belongs to you, so "let the user export their project" and "delete a user's data" are features you build.
- Live keys need a plan with
api_access(Studio, Team or Enterprise), an accepted developer agreement and an app approved in review.
Right for: a SaaS product, a pipeline, an internal tool, anything where Vidmoat is an implementation detail.
Model 2: Sign in with Vidmoat (OAuth 2.0 + PKCE)#
Your users connect their own Vidmoat accounts. Full protocol details are on the Authentication page; the shape:
- Their plan, their credits, their export allowance, their content. You are not reselling anything.
- Scoped consent, shown to the user. Spending scopes appear in a separate highlighted block and the Allow button tells the user that allowing includes spending their credits. Expect a materially lower approval rate when you ask for them; ask at the moment they are needed.
- For a client that belongs to a developer app, access tokens last 1 hour and refresh tokens 90 days, rotating on every use. OAuth tokens are always live, never test.
- Revocable, with no callback. The user revokes at
https://www.vidmoat.com/oauth/connectedand every token for your app is deleted immediately.invalid_granton refresh is not retryable: it means "this user disconnected us". scopeis required on the authorize call; there is no default. Store the granted scope from the token response, not the one you requested, because it may have been clamped to your app's ceiling.- Read
https://api.vidmoat.com/.well-known/oauth-authorization-serverrather than hardcoding endpoints. - Do not use
/api/oauth/register(Dynamic Client Registration). It exists so MCP clients can self-register; it produces a client with no app identity, no branding and no entry on the user's connected-apps page.
Right for: a tool that augments Vidmoat, a marketplace integration, anything where the user already has (or should have) a Vidmoat account.
The honest trade-off#
Model 1 is simpler to ship, gives you total control of the experience, and puts the entire cost on you. Model 2 shifts the cost to the user and gives them a revoke button, at the price of a consent screen in your funnel, refresh-token lifecycle code, and a permanent dependency on a relationship you do not own.
Pick Model 1 if your product is the video. Pick Model 2 if your product is about the user's Vidmoat account.
Scopes are enforced#
Eighteen scopes, in full:
| Scope | Grants | Spends |
|---|---|---|
account.read | /v1/me, /v1/usage, GET /v1/plugins | |
projects.read | read projects and workspaces; /validate | |
projects.write | create, /commands, PATCH, DELETE, workspaces | |
media.read | /v1/media | |
media.write | /v1/media/uploads, /v1/media/import | |
render.read | render status, /preview (all three formats) | |
render.write | POST /v1/renders | ✅ export allowance |
ai.transcribe | /v1/ai/transcriptions | ✅ |
ai.speech | /v1/ai/speech | ✅ |
ai.image | /v1/ai/images, /v1/ai/stickers | ✅ |
ai.video | /v1/ai/videos, /v1/ai/jobs/{id} | ✅ |
ai.agent | /v1/ai/agent | ✅ |
ai.analyze | (no endpoint ships yet) | |
stock.read | /v1/stock/search | |
webhooks.manage | (no v1 endpoint; webhooks are managed in the developer console) | |
plugins.invoke | POST /v1/plugins/{slug}/{tool} (sends the call's arguments to a third party) | |
telegram.users.read | GET /v1/telegram/users, GET /v1/telegram/users/{user} | |
telegram.users.write | PATCH /v1/telegram/users/{user} (live keys only) |
Keys minted before scopes existed carry every scope except webhooks.manage, plugins.invoke, telegram.users.read and telegram.users.write.
A missing scope is a 403 that names it:
Nothing is leaked by that (the vocabulary is public and the caller already holds the credential), and it is the difference between a ten-second fix and a support thread.
A scope is necessary but never sufficient. Holding ai.video does not mean the call succeeds. The owner's plan must include ai_generative_media and they must have credits, or you get a 402. This is deliberate: without it, minting a key with a scope would be a way to buy Studio features for free.
403 is your bug. 402 is the account's balance. Do not retry either.
7. What it costs#
The credit table#
Every action that spends credits, and what it spends:
| Action | Credits | Endpoint |
|---|---|---|
agent_message | 5 | POST /v1/ai/agent |
auto_caption | 10 per 5 minutes of media, rounded up | POST /v1/ai/transcriptions (src mode) |
ai_sticker | 20 | POST /v1/ai/stickers |
ai_image_gen | 60 on Grok; OpenAI 128 at 1k or 256 at 2k, plus 8 per reference; Vertex 128 at 1k or 192 at 2k, plus 2 per reference (up to 5 references) | POST /v1/ai/images |
ai_tts | 15 on Grok, 30 on OpenAI (text over 1,000 chars is truncated and reported, not rejected) | POST /v1/ai/speech |
ai_video_gen | billed per second (see below) | POST /v1/ai/videos |
agent_step (3), plan_workflow (3) and browser_record (40) exist in the table but no v1 endpoint charges them.
Video bills per second#
This is the one that will surprise you if you read the old table price of 500.
Choose with provider (grok, the default, qwen or wan). Through this API Grok always uses grok-imagine-video; the 1.5 model is not selectable.
So on grok-imagine-video: 5 s (the default) = 400 credits, 10 s = 800, 15 s = 1,200. The flat 500 was really a 15-second price that short clips overpaid for; per-second billing means margin does not depend on the caller's choice, and short clips got cheaper.
durationSec is a whole number of seconds, 5 by default. Grok takes 1 to 15 (out-of-range values are clamped), HappyHorse 3 to 15 and Wan 2 to 15. The same number is both billed and sent to the provider, so you cannot ask for 900 and be billed for 15 while getting 15, nor ask for 0.4 and be billed nothing.
The response tells you the basis so you can show a user a bill they can check:
Video is asynchronous (202) and one at a time per account: a second concurrent request is a 402 quota_exceeded, deliberately, because a caller looping video generation can spend a Studio allowance in under a minute. GET /v1/ai/jobs/{id} (scope ai.video) polls only video jobs; images, speech and stickers are synchronous. It polls through to the provider, persists the transition, and refunds on upstream failure exactly once (compare-and-swap on the job's own params, so two concurrent polls of a just-failed job cannot both pay out). A network blip talking to the provider comes back as still-PROCESSING with a warning, never as FAILED. Job statuses are PROCESSING, COMPLETED and FAILED.
Transcription#
- Free mode:
{ script, duration, start? }distributes text you already have across a span. No provider, no credits, and the output is the samewords[]shape. For programmatic use (a voiceover you wrote, a script you are narrating) this is almost always the right call. - ASR mode:
{ src, language? }. 10 credits per 5 minutes, rounded up. So a 12-minute podcast is 3 blocks = 30 credits. - Hard cap: 30 minutes (
MAX_CAPTION_SECONDS = 1800). Longer is a400telling you to split it and offset each part's word times by its start. - The length is probed before you are charged. Unreadable length is a
400, not a guess: an unbounded bill and a genuinely broken source both deserve an error. - If the deployment has no ASR provider, you get
501 not_configuredbefore any charge.
Refunds#
Every generation endpoint charges before the provider call and refunds on failure (charge() hands back its own refund, so the amount given back cannot drift from the amount taken). Charging afterwards would let a caller disconnect mid-request and get the work free.
Test-mode keys spend nothing#
A vmk_test_… key returns deterministic fixtures. Ids are a hash of the request, so POST /v1/renders twice with the same arguments returns the same job id, and GET /v1/renders/{that id} answers from the id with no row ever existing, which is what lets you exercise your poll loop.
- Fixtures come from
public/fixtures/v1/and are real, valid media (render.mp4|webm|gif,image.jpg,video.mp4,speech.mp3,sticker.png) served withACAO: *, so they decode in a<video>element on your origin. - Test renders come back
COMPLETEDimmediately with a sample fixture andtest: true. They are never queued, use no export quota, and write no row. - Generation returns sample fixtures and spends nothing;
quoteOnlyis ignored on a test key. - The agent still needs a real
projectId, then returns an empty plan without calling a model. - Plugins are not called, and
PATCH /v1/telegram/users/{user}is refused. - Validation still runs first: a bad
projectIdstill 404s, an empty project still 400s. A sandbox that accepts anything teaches nothing. - Test mode is not a full sandbox. Only the paths that cost money or occupy a worker are stubbed. Projects,
/commands,/media/uploads,/media/importand previews do real work with a test key: real rows, real files. Plan accordingly. - A test-mode request that ever reaches the billing path throws a loud
500rather than silently spending. That is a tripwire for a missing branch, not a path you should see.
Rate limits and concurrency#
| Hobby | Creator | Studio | |
|---|---|---|---|
| Requests/min (per key) | 20 | 120 | 600 |
| Monthly credits | 40 | 1,500 | 6,000 |
| Concurrent renders | 1 | 3 | 5 |
| Monthly exports | 8 | unlimited | unlimited |
| Max export height | 720 | 2160 | 2160 |
| Max upload MB (plan cap) | 200 | 2,048 | 8,192 |
| Max projects | 3 | unlimited | unlimited |
| Live API keys | no | no | ✓ (api_access) |
- Team and Enterprise keys get 600 requests/min, the same as Studio (the higher of the owner's own plan and their team's plan applies).
- The limiter is keyed on the key id, not the user, over a 60-second sliding window, so two of your integrations cannot starve each other and a leaked key can be throttled by revoking it.
- Frame previews have their own extra bucket: 10/min per key.
PATCH /v1/telegram/users/{user}is limited to 120/min per app.- An app the platform has marked
throttleddrops to 6/min and keeps working; that is the whole point ofthrottledversussuspended. - Every response that reaches the limiter carries
X-RateLimit-Limit,X-RateLimit-RemainingandX-RateLimit-Reset.X-RateLimit-Resetis the window length in seconds (normally 60), not a Unix timestamp. A429or503also carriesRetry-After. Refused requests do not count against the window. - If the limiter's store (Redis) is down, v1 fails closed: every request gets
503 unavailablewithRetry-After: 30, and nothing is run or charged. - Inline uploads are capped at 32 MB regardless of plan (the path buffers in memory); the plan cap applies on top. Bigger files go via
POST /v1/media/import, up to your plan's upload cap or 512 MB, whichever is smaller.
8. Production concerns#
Idempotency: there isn't any, so build it#
There is no Idempotency-Key header. Commands are not idempotent. What the platform does give you:
- Uploads dedupe by sha256. Re-POSTing the same bytes returns the existing file with
deduplicated: true, so upload retries are safe and do not multiply your disk bill. - Test fixture ids are deterministic, so retrying a test call is stable.
- Video refunds are claimed with a compare-and-swap, so concurrent polls cannot double-refund.
Everything else is on you. The pattern that works:
For "create a project and edit it" specifically, use the commands array on POST /v1/projects: one call, one failure domain, and a partial failure leaves you a project plus a results array instead of an ambiguous half-state.
Polling cadence#
The limiter does not distinguish an eager poller from an attack. Both are requests against the same per-minute budget.
- Renders: 3 to 5 seconds is right. A
PENDINGrender that has not started will not start faster because you asked twice. - Video generation: 5 to 10 seconds. It takes tens of seconds at best.
- Budget it: on Creator (120/min), one client polling every 2 seconds burns 30 of your 120. Ten concurrent jobs at that cadence is 300/min and you are throttled.
- Back off on
429, honourRetry-After, and stop polling terminal states:COMPLETED,FAILEDand (for renders)CANCELLEDanswer from the stored row, so re-polling is pure waste. - A transient
warningon a still-PROCESSINGgeneration job means keep going, not give up.
402 handling#
Three distinct codes share the 402 status, and they need three different behaviours:
| Code | Means | Do |
|---|---|---|
plan_required | the owner's plan lacks the feature | upsell; never retry |
insufficient_credits | balance too low (carries remaining) | tell the user; never retry |
quota_exceeded | a plan limit: projects, exports, concurrency, or the one-video-at-a-time fairness rule | for concurrency, wait and retry; for exports/projects, upsell |
The only retryable one is concurrency, and even then with backoff.
The error envelope#
code is stable and machine-readable; message is prose written for a human and will be reworded. Branch on code. Never parse message.
| Code | Status | Meaning |
|---|---|---|
invalid_request | 400 (409) | malformed body or params; a 409 is a project revision conflict and carries conflictCode (stale_rev or help_handoff_locked) and rev |
unauthorized | 401 | no credential, or unknown/revoked/expired |
insufficient_scope | 403 | valid credential, wrong permissions (carries scope, granted) |
access_suspended | 403 | the owner's account is suspended (carries restriction) |
app_suspended | 403 | this app's kill switch, owner is fine |
plan_required | 402 | plan lacks the feature |
insufficient_credits | 402 | carries remaining |
quota_exceeded | 402 (413) | a plan limit; 413 when storage is full |
not_found | 404 | no such resource for this owner |
payload_too_large | 413 | over the upload cap |
prompt_rejected | 422 | refused by prompt safety (carries field) |
rate_limited | 429 | honour Retry-After |
provider_error | 502 | an upstream failed; retrying is correct |
not_configured | 501 | capability not enabled on this deployment |
internal_error | 500 | ours |
unavailable | 503 | a dependency (the rate limiter) is briefly down; honour Retry-After, nothing was run or charged |
Codes are added, never renamed. Treat an unknown code as a generic failure of its HTTP class.
Note that someone else's project id is a 404, not a 403: distinguishing them would confirm the id exists.
Vidmoat-Version is echoed on every response (default 2026-08-01). Nothing branches on it yet; send it anyway, so the day some behaviour becomes date-pinned your clients are already sending a version they can see acknowledged.
You are responsible for what your users generate#
Third-party apps mean content generated at Vidmoat's expense, hosted on Vidmoat's domain, by a user Vidmoat has no relationship with. So:
- Prompt filtering runs server-side, in the shared wrapper rather than per-endpoint, on every generation route. Any of
prompt,text,message,script,captionin the body is checked before any charge. A refusal is422 prompt_rejectedand costs nothing. - Your app can be suspended independently of your account.
DeveloperApp.statusisactive | throttled | suspended;throttleddrops you to 6 requests/minute,suspendedreturns403 app_suspendedon every call. One bad integration is stoppable without banning the person who wrote it. - A monitored contact address, an accepted developer agreement and an app approved in review are required before a live key is issued. That is what makes takedown possible.
- Everything is attributed to the owner:
UsageEventandApiRequestLogcarryappIdandkeyId, andRenderJobdoes too. There is always a real Vidmoat user behind every byte. - Do your own filtering in front of Vidmoat's. Server-side prompt safety is a backstop, not your moderation policy, and "the API let it through" is not a defence.
Observability#
GET /v1/usage?days=30 gives two breakdowns, and the second is the one you want:
account includes the owner's own editing in the browser. app covers only what came through this app's credentials, which is the number you need when you get a credit bill. credits is net, so refunds reduce it rather than reading as a second charge. days caps at 90.
9. Worked example: a podcast episode into five captioned vertical clips#
This exercises transcription, the command vocabulary, per-clip projects, preview verification and rendering. It is roughly the smallest thing that is a real product.
Inputs: a URL to a 40-minute episode (audio or video), and a way to pick interesting spans (your model, your heuristic, or a human).
Scopes: media.write projects.write projects.read render.write render.read ai.transcribe.
Cost: the episode is over the 30-minute transcription cap, so it must be split. Two halves of 20 minutes = 4 blocks each = 40 credits each = 80 credits total. Then 5 renders against the export allowance. No generative credits at all.
What this example is teaching, beyond the code#
- Transcription has no time-range parameter. Chunking is your job, and so is re-offsetting word times. The 30-minute cap is not negotiable and it is enforced from a probe before you are charged, so you find out cheaply.
addCaptionstakes timeline times. It derives each caption clip'sstartfrom its line's first word and stores per-word offsets relative to that. Feed it raw episode times and every caption lands 20 minutes into a 40-second clip.- Positions are pixels from the canvas centre, +y down. Getting this wrong is the single most common way an otherwise-correct program produces a wrong frame, which is why
lintexists and why you should look at one frame per clip. - Read
results[].data.clipId. You cannot know a clip's id before you create it, and collision handling may not even have put it on the track you asked for. - Serialise your renders. Concurrency is 1/3/5 by plan.
- Free credits where you can get them. If your product generates the narration itself, use
{ script, duration }transcription mode: it costs nothing, needs no provider, and returns the samewords[].
Extensions that are one command each#
Once the pipeline above works, most of what a competitor charges for is a single op:
- Beat-cut b-roll:
detectBeatsclient-side (§4.2) →addMarkers→sliceClip times=[…]. - Cut filler words: the transcript already has word timings →
cutRangeswithripple: true, which closes the gaps and pulls later clips left. - Music under speech:
addClipthe bed, thenautoDuck. One command rewrites the music's volume keyframes to dip wherever voice plays, with smooth ramps. A flat music bed over dialogue is the tell of an amateur edit. - One global look:
addAdjustmentLayer, thensetColoron the returned clip id. Grades everything below it instead of every clip individually. - Trim dead air:
detectSilenceclient-side →trimSilence. - A designed title card:
addHtmlElementwith your own HTML/CSS. System fonts, gradients, inline SVG and emoji work; scripts and external resources do not (it rasterises through an SVGforeignObject).
Appendix A: every shipped endpoint#
Everything that is live under /v1 today, and nothing that is not.
| Method | Path | Scope | Cost |
|---|---|---|---|
| GET | /v1/me | account.read | none |
| GET | /v1/usage?days= | account.read | none |
| GET | /v1/schema/commands | (none; credential still required) | none |
| GET | /v1/projects?limit&cursor&view | projects.read | none |
| POST | /v1/projects | projects.write | project quota |
| GET | /v1/projects/{id}?view= | projects.read | none |
| PATCH | /v1/projects/{id} | projects.write | none |
| DELETE | /v1/projects/{id} | projects.write | none |
| PUT | /v1/projects/{id}/workspace | projects.write | none |
| GET | /v1/workspaces | projects.read | none |
| POST | /v1/workspaces | projects.write | none |
| PATCH | /v1/workspaces/{id} | projects.write | none |
| DELETE | /v1/workspaces/{id} | projects.write | none |
| POST | /v1/projects/{id}/commands | projects.write | plan-gated ops dropped |
| POST | /v1/projects/{id}/validate | projects.read | none |
| GET | /v1/projects/{id}/preview?format&at&resolution | render.read | none (10 frames/min) |
| POST | /v1/renders | render.write | export quota |
| GET | /v1/renders?limit&cursor&projectId&status | render.read | none |
| GET | /v1/renders/{id} | render.read | none |
| GET | /v1/media?limit&cursor | media.read | none |
| POST | /v1/media/uploads (multipart, file) | media.write | upload cap |
| POST | /v1/media/import | media.write | upload cap |
| GET | /v1/stock/search?query&type&limit | stock.read | none |
| POST | /v1/ai/transcriptions | ai.transcribe | 0 or 10/5min |
| POST | /v1/ai/speech | ai.speech | 15 (Grok) or 30 (OpenAI) |
| POST | /v1/ai/images | ai.image | 60 (Grok) |
| POST | /v1/ai/stickers | ai.image | 20 |
| POST | /v1/ai/videos | ai.video | per second, from 80/sec (Grok) |
| GET | /v1/ai/jobs/{id} | ai.video | none |
| POST | /v1/ai/agent | ai.agent | 5 |
| GET | /v1/plugins | account.read | none |
| POST | /v1/plugins/{slug}/{tool} | plugins.invoke | none |
| GET | /v1/telegram/users?q&sort&page | telegram.users.read | none |
| GET | /v1/telegram/users/{user} | telegram.users.read | none |
| PATCH | /v1/telegram/users/{user} | telegram.users.write | none (120/min per app) |
Not shipped: /v1/ai/analyze, /v1/webhooks. Do not design against either.
Deliberately excluded, permanently or for now: browser recorder (holds a CPU core for ~3 minutes on the box serving the API), inpainting (one global CPU worker capped at 1/user), Flows (an orchestration product, not a primitive), Vidmoat Social (a moderation liability with no revenue; no scope grants it), channels, CDS, admin, billing (internal surfaces with no versioning promise).
Cursor pagination: GET /v1/projects, /v1/renders and /v1/media all return { data, hasMore, nextCursor }. Cursors are opaque base64; do not parse them. limit defaults to 20 and caps at 100.
Appendix B: the versioning contract#
Additive changes ship into v1: new fields, new endpoints, new enum values. Your client must ignore unknown fields and tolerate unknown enum values and error codes. That is the promise you are making in exchange for the one Vidmoat is making, which is that no field already present in a v1 response is removed or repurposed. Breaking changes create /v2, with v1 supported for at least twelve months.
Send Vidmoat-Version: 2026-08-01. It is reserved for date-pinned behaviour inside v1 and is echoed back today.
Questions: developers@vidmoat.com. The portal, keys and interactive docs are at https://developer.vidmoat.com.