Skip to content

16 models · 4 modalities

Choose the model for the shot, not the whole project.

Storyloom keeps a curated provider catalog behind one typed action surface. Explore fast, compare deliberately, and use the model that fits each part of the sequence.

A continuous signal passing through varied image, video, audio, and 3D forms.

6 in source

Image models

The cards below render directly from the same catalog the Storyloom model browser uses.

Nano Banana Pro

Google's frontier image model — strong prompt adherence and text rendering.

medium latency$0.15 / image

FLUX1.1 [pro] ultra

High-res, photoreal image generation up to 4MP.

medium latency$0.06 / image

FLUX.1 [dev]

Fast, high-quality 12B open-weight image model.

fast latency$0.03 / image

GPT Image 1

OpenAI's image model — excellent instruction following and typography.

slow latency$0.04 / image

Recraft V3

Design-grade images with brand styles, vector art, and long text.

medium latency$0.04 / image

Ideogram V3

Best-in-class text-in-image and graphic design layouts.

medium latency$0.06 / image

5 in source

Video models

The cards below render directly from the same catalog the Storyloom model browser uses.

Veo 3

Google's flagship text-to-video with native audio.

slow latency$0.75 / sec

Kling 2.0 Master

Cinematic motion and strong prompt adherence, text- or image-to-video.

slow latency$0.28 / sec

Luma Ray 2

Fluid, coherent motion with realistic physics.

slow latency$0.10 / sec

Hailuo 02

MiniMax text-to-video with lively, expressive motion.

slow latency$0.04 / sec

Wan 2.2 A14B

Open-weight text-to-video with high motion quality.

medium latency$0.08 / sec

4 in source

Audio models

The cards below render directly from the same catalog the Storyloom model browser uses.

ElevenLabs TTS (via fal)

Multilingual, natural text-to-speech.

fast latency$0.10 / 1K chars

Stable Audio

Text-to-audio for music beds and sound design.

medium latency~$0.02 / gen

MMAudio V2

Generate synchronized sound effects for a video clip.

fast latency$0.001 / sec

F5 TTS

Fast open-weight voice cloning text-to-speech.

fast latency$0.05 / 1K chars

1 in source

3D models

The cards below render directly from the same catalog the Storyloom model browser uses.

Hyper3D Rodin

Text-to-3D — generates a rotatable GLB mesh with PBR materials.

slow latency$0.40 / gen

Bring your own key

Connect the provider account you already control.

Keys added in the studio are stored in your browser, forwarded to your Storyloom deployment for the request, and then used with the selected provider. No database stores them.

  1. 01

    fal.ai

    Fast serverless image, video & audio models (FLUX, Veo, Kling, Ray, and more).

  2. 02

    ElevenLabs

    Text-to-speech, voice cloning, music & sound effects.

  3. 03

    HeyGen

    AI avatar & talking-head video generation from a script.

Open Extensions in the studio

Honest estimates

See the unit before you run the action.

Catalog prices are stored alongside each model and used for preflight estimates. Per-second and per-character actions depend on the requested output, so provider invoices remain the source of truth.

Model catalog contract
type Modality = 'image' | 'video' | 'audio' | 'model'

interface FalModel {
  id: string
  kind: Modality
  costPerUnit: number
  unit: 'image' | 'second' | 'kchar' | 'generation'
  latencyHint: 'fast' | 'medium' | 'slow'
}

The right tool for each frame

Keep model choice flexible and the creative system intact.

Open Storyloom without an account or a provider key. Add your own models when the direction is ready for production.

Models · Storyloom