Skip to content

Video preview

Play the app compositor while editing, and inspect exact frames before exporting. None of it uses your export allowance.

Embed the app player#

The /embed/player page mounts the same preview canvas as the web editor and the phone app. It plays video, audio, titles, effects and keyframes without exporting. The editor starter includes a working React integration.

Fetch the document with GET /v1/projects/{id}?view=document on your server, keep the key there, and send only the document to the player, with absolute, browser-accessible media URLs.

player.js
const origin = 'https://www.vidmoat.com';
iframe.src = origin + '/embed/player?parentOrigin=' +
  encodeURIComponent(location.origin);
iframe.allow = 'autoplay; fullscreen';
const channel = 'vidmoat.player.v1';
const send = data => iframe.contentWindow.postMessage({ channel, ...data }, origin);
window.addEventListener('message', event => {
  if (event.source !== iframe.contentWindow || event.origin !== origin ||
      event.data?.channel !== channel) return;
  if (event.data.type === 'ready') {
    send({ type: 'document', document: project.document, revision: 1, time: 0 });
  }
  // 'loaded' acknowledges the document revision.
  // 'state' reports time and playing; 'error' includes a message.
});
// After 'loaded', from your playback controls:
send({ type: 'play' });
send({ type: 'pause' });
send({ type: 'seek', time: 2.5 });
// After an edit, send 'document' again with a new revision.

Fetch a single frame#

The preview endpoint draws one timeline timestamp through the same pipeline that renders the export: clips, keyframes, effects, captions, colour grade. What you see is what the file will contain at that moment.

Shell
curl "https://api.vidmoat.com/v1/projects/$ID/preview?format=image&at=2.5" \
  -H "Authorization: Bearer $VIDMOAT_KEY" \
  -o frame.jpg

It needs render.read, not render.write: a preview uses no export allowance, so a read-only key can check a composition it is not allowed to render. Frames are limited to 10 a minute per key.

QueryMeaning
formatjson (default): metadata, lint and suggestedPreviewTimes, and a frame as a data URI when at is set. image: the JPEG bytes. html: the composition below.
atTimeline time in seconds. Defaults to 0 for image.
resolutionfull or half, for html only. Frames are captured at half the project resolution, which is enough to check layout.

Lightweight HTML composition (approximate)#

format=html returns an approximate HTML version of the timeline: clips, text, captions, effects and keyframes, driven by a single GSAP timeline. No file is produced. Serve it in an iframe and you have a scrubbable preview.

It arrives paused at 0. Clip visibility is driven by the timeline, so until you tell the document what time it is, nothing is on screen. That is what makes deterministic seeking possible, but an un-driven composition looks like a black rectangle.

Live: this is format=html, not a video. Built from the commands below and served with no key; open the raw document to read exactly what you get back.

The commands that produced it. Note where fontSize and y live: passing them at the top level is ignored, and every element lands centred on top of everything else.

JSON
{
  "commands": [
    {
      "op": "addClip",
      "type": "video",
      "src": "/fixtures/v1/video.mp4",
      "start": 0,
      "duration": 6,
      "trackIndex": 0
    },
    {
      "op": "addTextClip",
      "text": "Shipped from the API",
      "start": 0.4,
      "duration": 3,
      "trackIndex": 1,
      "style": {
        "fontSize": 72
      },
      "patch": {
        "y": -220
      }
    },
    {
      "op": "addTextClip",
      "text": "No editor. No timeline. One POST.",
      "start": 1.2,
      "duration": 4.2,
      "trackIndex": 1,
      "style": {
        "fontSize": 34
      },
      "patch": {
        "y": -110
      }
    },
    {
      "op": "addCaptions",
      "words": [
        {
          "text": "Every",
          "start": 3.4,
          "end": 3.7
        },
        {
          "text": "frame",
          "start": 3.7,
          "end": 4
        },
        {
          "text": "you",
          "start": 4,
          "end": 4.2
        },
        {
          "text": "see",
          "start": 4.2,
          "end": 4.45
        },
        {
          "text": "here",
          "start": 4.45,
          "end": 4.75
        },
        {
          "text": "is",
          "start": 4.75,
          "end": 4.9
        },
        {
          "text": "real",
          "start": 4.9,
          "end": 5.4
        }
      ],
      "trackIndex": 2
    }
  ]
}

Drive it from the parent frame. Both handles are on the composition's own window:

JavaScript
const win = document.querySelector('iframe').contentWindow;

win.setCompositionTime(2.5);  // seek, and pin every video to its own in-clip time
win.playComposition();        // or let it run
win.pauseComposition();       // pause picture and sound together

Use setCompositionTime rather than seeking the GSAP timeline directly. tl.seek() suppresses events by default, so the callback that decides which clips are on screen never fires, and GSAP does not know a <video> has a playhead of its own.

What the frames look like#

The same composition at four timestamps. They are the frames worth sampling in any project: where something appears, where two elements share the screen, where one track hands over to another, and the last frame before a clip ends.

The demo composition at 0.8 seconds
at=0.8 Title only. Check it clears the top safe area.
The demo composition at 2 seconds
at=2 Title and subtitle together: the overlap moment.
The demo composition at 4.3 seconds
at=4.3 Captions running, titles gone.
The demo composition at 5.8 seconds
at=5.8 Last frame before the clip ends.

Preview inline with the edit (MCP)#

If an agent drives Vidmoat over MCP, do not preview as a separate step. Pass previewAt on the edit itself and the frames come back as images in the same response:

JSON
{
  "name": "edit_project",
  "arguments": {
    "projectId": "…",
    "commands": [{ "op": "addTextClip", "text": "Hello", "start": 0, "duration": 3 }],
    "previewAt": [0.5, 1.5, 2.5]
  }
}

Up to three timestamps. The standalone preview_frame tool takes projectId and time, for looking without editing.

Which one to use#

You areUse
A server-side pipeline checking its own outputGET /v1/projects/{id}/preview
Building a signed-in web UI on top of VidmoatThe app player at /embed/player
An agent making edits and needing to see themMCP previewAt on the edit call
Shipping a finished video to a userA real render. Previews are not deliverables.