# Eleven Labs Financial Model

AI-powered automatic dubbing SaaS that converts audio/video content into any language while preserving the original speaker's voice, emotion, and intonation.

- Canonical: https://finamodel.com/startups/eleven-labs
- Excel download: https://finamodel.com/startup-models/eleven-labs.xlsx
- Category: AI/ML
- Model type: SaaS ARR / Valuation
- Funding round: Pre-Seed
- Funding: $2M
- Founded: 2023
- Geography: Global (content creators worldwide); headquartered UK (Eleven Labs Ltd. per copyright notice).
- Customer: B2B

## About the company

ElevenLabs automates dubbing by converting uploaded audio or video into another language while retaining the speaker's voice, emotion, and intonation. Its Speech-plus-Text approach supports voice cloning and can produce a dubbed ten-minute video in roughly two minutes.

The SaaS product combines a subscription bundle of convertible minutes with per-minute overage pricing, illustrated at about $1 per minute. It targets creators first, with potential API and SDK distribution to editing software and game developers, plus a human-in-the-loop quality option.

The model cohorts paying subscribers and translates videos, minutes, and languages into subscription and overage revenue. It tests creator acquisition, usage, churn, inference and storage cost, gross margin, API adoption, product and marketing headcount, cash burn, and runway.

## What's included

- 5-year monthly revenue build with stage-appropriate growth assumptions
- Full P&L, headcount plan, and operating-expense schedule
- Cash-flow statement, runway, and burn-rate tracking
- Valuation via exit multiple with a DCF cross-check
- Returns analysis with MOIC and IRR
- Unit economics including CAC, LTV, payback, and cohort retention

## Product & value proposition

- Automated dubbing SaaS: upload audio or video, click to dub into another language, download result in ~2 minutes for a 10-minute video.
- Core differentiator: novel speech representation that takes both Speech and Text as inputs (vs. traditional TTS), preserving prosody (per-phoneme, speaker-independent annotations) and speaker voice (separate speaker embeddings).
- Three headline features:
  1. Human Quality - preserves emotions, intonation, performance via model trained on thousands of hours of professional dubbing.
  2. Personalised - voice cloning across languages (speaker's own voice maintained).
  3. Simple & Quick - end-to-end SaaS; human-in-the-loop option for quality uplift.
- Traditional dubbing benchmark: ~$100/min and >2 weeks per 10-minute video.

## Market

Three-tier expanding market framing:
- $2B/yr - TAM for all professional content creators (podcasts + videos)
- $4.6B/yr - Current spend on game localization + movie dubbing (existing industry to disrupt)
- $24B/yr - Full localization, translation, interpreting total market

Content creator funnel deep-dive:
- 50M+ content creators worldwide (TAM)
- 2M professional creators (SAM)
- 100K YouTube creators with >500K subscribers (SOM)
- 10K creators that upload captions (Immediate Market)
- Content volume from SOM: 9M minutes/month (3 videos × 10 min × 3 languages × 100K creators)
- Implied revenue at ~$1/min: $110M/year

No market growth rate cited in deck.

## Revenue model

- Primary model: subscription with a base bundle of convertible minutes.
- Usage overage: ~$1/min of dubbed audio (company's own illustration of floor price; actual pricing not disclosed).
- Channels: direct SaaS (self-serve), potential API/SDK for audio & video editing software and game developers.
- Adjacent use-cases on roadmap (not yet monetised): real-time dubbing, real-time voice conversion, professional dubbing for film, localization/advertising, offline voice generation (games, audiobooks, podcasts).

## Traction & metrics

- Traction & Feedback slide (slide 11): fully redacted - no revenue, customer count, MRR, or growth metrics available.
- MrBeast case study used as proxy demand signal: Spanish dubbed channel (launched 2021) reached 19M subscribers; one dubbed video generates ~$50K for the creator. (This is third-party data, not Eleven's own traction.)
- Prototype exists with stated demo dub time of 2 minutes for a 10-minute video.

## Competition / moat

Competitive quadrant (Quality vs. Accessibility & Speed):
- Top-left (high quality, low speed): traditional studios - Logic Union, Simonsays(?), MN Creative - voice actors, post-production.
- Centre (semi-automated, high manual intervention): Deepdub.ai, Papercup.
- Bottom-right (accessible/fast, lower quality): Amazon Polly, IBM Watson, Google Wavenet - pure TTS.
- Top-right (high quality + high speed): Eleven Labs - positioned as sole occupant.

Moat claims:
1. Novel Speech+Text input architecture (not pure TTS) - prosody mapping preserves emotional performance across languages.
2. Data flywheel: high-volume creator usage improves speech/text datasets.
3. Trained on thousands of hours of professional dubbing data.
4. Generalizable to new languages quickly.

## Team & funding ask / use of funds

Team:
- Mati Staniszewski, CEO - Mathematics, Imperial College London; ex-Palantir (deployment); ex-BlackRock, Opera Software; founded Mathscon (1,000+ students).
- Piotr Dabkowski, CTO - CS, Cambridge & Oxford; ex-Google ML; NeurIPS paper (300+ citations); created Js2Py (250K+ downloads/month).
- Both friends since high school, studied and worked together.

---

## Recommended financial model

**Archetype + why:** Usage-based SaaS with subscription base + per-minute overage - mirrors the company's own stated model. Build as a bottoms-up consumption model: cohorted subscriber acquisition driving monthly active dubbing minutes, layered with a subscription floor and variable per-minute revenue above the bundle. This is analogous to a cloud-infra usage model (think Twilio/AWS), not a pure seat-based SaaS.

**Forecast horizon & granularity:** Monthly for Years 1–2, quarterly for Years 3–5. Five-year total horizon is appropriate for a seed-stage AI SaaS.

**Key drivers & assumptions:**

| Driver | Value | Source |
| -- | -- | -- |
| Content creator TAM | 50M worldwide | - |
| Professional creator SAM | 2M | - |
| SOM (YT >500K subs) | 100K | - |
| Avg videos/creator/month | 3 | - |
| Avg video length | 10 min | - |
| Languages dubbed per video | 3 | - |
| Floor price per minute | ~$1/min audio | - |
| Initial market (caption uploaders) | 10K creators | - |
| Monthly minutes per creator (SOM calc) | 90 min/mo | - |
| Traditional dubbing cost | ~$100/min | - |
| Traditional dubbing lead time | >2 weeks per 10-min video | - |
| Month 1 paying subscribers | 500 | Seed stage; prototype live; redacted traction suggests some early users |
| Monthly subscriber growth rate (Y1) | 15–25% MoM | Typical viral/product-led SaaS at early stage with strong creator word-of-mouth |
| Avg monthly minutes consumed per subscriber | 90 min/mo | Based on deck's own 3 videos × 10 min × 3 languages assumption |
| Subscription ARPU (base plan) | $30–$99/mo incl. bundle | Benchmarked against Descript, Otter.ai, Riverside; deck implies ~$1/min floor |
| Overage rate | $1.00/min above bundle | Deck's stated floor price |
| Bundle minutes included in subscription | 60 min/mo | Encourages overage; typical usage model |
| Gross margin (long-run) | 60–70% | AI inference + storage costs; speech model compute is material at early scale, improves with volume |
| Gross margin (Year 1) | 40–50% | Compute-heavy early; GPU cost per inference is high before scale economies |
| Monthly churn | 4–6% | Creator-segment tools see moderate churn; no retention data in deck |
| CAC | $50–$150 per subscriber | Product-led / self-serve with content creator influencer channel; no data in deck |
| S&M as % of revenue (Year 1) | 30% | Early land-and-expand; low touch |
| R&D as % of revenue (Year 1) | 40% | Two founders, ML-heavy; compute + talent |
| G&A | $20K–$40K/mo | Seed stage, UK entity |

**Scenarios (Base / Bull / Bear - which variables flex):**
- Bear: Subscriber growth 10% MoM, churn 7%, gross margin stays at 40% (GPU costs don't scale down), ARPU at $30. Competitors close the quality gap faster.
- Base: Subscriber growth 20% MoM, churn 5%, gross margin reaches 65% by Year 3, ARPU $60 blended.
- Bull: Subscriber growth 30% MoM driven by viral creator adoption (MrBeast-style flywheel), churn 3%, expansion into professional film/game studio segment lifts ARPU to $150+, gross margin 70%.

**Required sheets / outputs:**
1. Assumptions - all drivers in one editable block, colour-coded vs.
2. Subscriber Model - monthly cohort build: new adds, churn, net adds, ending subscribers by plan tier.
3. Revenue Model - subscription revenue (ARPU × subscribers) + overage (minutes above bundle × overage rate).
4. P&L (Income Statement) - Revenue → Gross Profit → EBITDA → Net Income; monthly Y1–Y2, quarterly Y3–Y5.
5. Headcount Plan - founders + ML engineers + sales/marketing hires; drives OpEx.
6. Cash Flow & Runway - burn rate, cash balance, months of runway; critical for seed-stage investors.
7. Market Sizing Summary - TAM/SAM/SOM waterfall tied to deck figures.
8. Scenario Toggle - dropdown (Bear/Base/Bull) that feeds growth rate, churn, ARPU, and margin assumptions.
9. Dashboard - KPI summary: MRR, subscribers, minutes dubbed, gross margin %, cash runway.

## Frequently asked questions

### Is the Eleven Labs financial model free?

Yes. The Eleven Labs model is a free Excel download with live formulas.

### Can I change the assumptions?

Yes. The workbook is editable and its live formulas recalculate when assumptions change.
