# Datagen Financial Model

Cloud-based self-service synthetic data generation platform for computer vision AI teams.

- Canonical: https://finamodel.com/startups/datagen
- Excel download: https://finamodel.com/startup-models/datagen.xlsx
- Category: AI/ML
- Model type: SaaS ARR / Valuation
- Funding round: Series B
- Funding: $50M
- Founded: 2022
- Geography: Israel-headquartered (academic advisors at Bar-Ilan, Technion); global customer base implied.
- Customer: B2B

## About the company

Datagen offers a cloud platform that generates labelled synthetic data for computer-vision models, removing manual collection, labelling, and post-processing. Its Faces and Humans in Context generators give users control over people, scenes, expressions, accessories, and multiple forms of machine-readable ground truth.

The platform addresses automotive, robotics, security, AR/VR, metaverse, industrial, and retail use cases. It sells through self-serve access and enterprise contracts to technology giants and Fortune 500 companies, using per-image or dataset volume pricing alongside platform subscriptions.

The model separates enterprise and self-serve cohorts, then prices generated volume by product line and modality. It tracks generator mix, compute cost, gross margin, large-account sales cycles, expansion-driven retention, simulation and R&D headcount, GPU investment, cash flow, and runway.

## What's included

- 5-year monthly revenue build with stage-appropriate growth assumptions
- Full P&L, headcount plan, and operating-expense schedule
- Cash-flow statement, runway, and burn-rate tracking
- Valuation via exit multiple with a DCF cross-check
- Returns analysis with MOIC and IRR
- Unit economics including CAC, LTV, payback, and cohort retention

## Product & value proposition

Cloud-based self-service platform that generates labeled synthetic data for computer vision models, eliminating manual collection, labelling, and post-processing. Key generators:
- **Faces Generator** - control over identity (age, gender, ethnicity), scene, gaze, expression, accessories; outputs RGB, infrared, depth, normal, semantic segmentation, 2D/3D keypoints.
- **Humans in Context (HIC)** - full-body humans in four domains: In-Cabin Automotive (DMS/OMS), Smart Office, Home Security, XR/Metaverse.
- **Objects in Context** - listed as "coming soon".

Value props: eliminates human error, bias, privacy risk; pixel-perfect 2D/3D ground truth; large-scale; fully privacy-compliant.

## Market

- No explicit TAM/SAM/SOM figures in deck.
- Market framing: Gartner prediction cited - "By 2030, Synthetic Data Will Completely Overshadow Real Data in AI Models". Area chart shows synthetic data share of AI training data rising steeply from 2020 to 2030, with real data share shrinking - no axis values given.
- Verticals addressed: Automotive, Robotics, Security, AR/VR, Metaverse, Industrial, Retail.
- Problem scope: 96% of organizations have problems with training data quality, quantity, and speed (Dimensional Research, May 2019).

## Revenue model

- Per-image / per-dataset volume pricing - common in synthetic data (e.g., price per 1,000 images generated); rationale: self-service + "large scale" framing.
- Subscription / seat tier for platform access - rationale: "cloud-based self-service" language typical of SaaS.
- Enterprise contract / custom generation for Fortune-500 accounts - rationale: "Tech Giants & Fortune-500 Companies" customer description implies negotiated deals.
- Channel: direct sales + self-serve cloud portal.

## Traction & metrics

- Founded: 2018.
- Team size: 85+ simulation experts.
- Customers: "Tech Giants & Fortune-500 Companies" - no count, no names, no revenue disclosed.
- No ARR, revenue, growth rate, churn, NPS, or dataset volumes disclosed in deck.

## Competition / moat

Not explicitly addressed in deck. Implied moats:
- Deep simulation expertise (85+ experts, world-class academic advisors: MPI-IS, BAIR, CMU, Technion, Bar-Ilan).
- Proprietary simulation engine capable of photorealistic humans + full scene control.
- Multi-modal ground truth output (RGB, IR, depth, normal, segmentation, keypoints) that real-data collection cannot match.
- Privacy-by-design (no real people → no GDPR/CCPA exposure).

## Team & funding ask / use of funds

- Academic advisors: Prof. Michael J. Black (MPI-IS), Prof. Trevor Darrell (BAIR), Prof. Gal Chechik (Bar-Ilan), Prof. Lihi Zelnik (Technion), Anthony Goldbloom (Kaggle CEO), Prof. Fernando De La Torre (CMU / FacioMetrics CEO).
- Founding team: Not named in deck (only advisors shown).

## Recommended financial model

- **Archetype + why:** Usage-based SaaS / data-as-a-service (DaaS) model, with an enterprise contract overlay. Revenue = volume of synthetic data generated (images or datasets) × price per unit, plus platform subscription seat fees. Rationale: product is self-service cloud generation at scale; buyers are large enterprises running recurring ML pipelines, so a blended usage + subscription model is standard. Not a marketplace (no third-party supply side) and not a simple ARR SaaS (usage variance too high to model as flat recurring).

- **Forecast horizon & granularity:** 5-year model (Year 1–5), monthly granularity for Year 1–2, quarterly for Year 3–5.

- **Key drivers & assumptions:**
  - New customers / quarter: start at 2–4 enterprise + 5–10 self-serve in Year 1; ramp ~30% YoY; rationale: Fortune-500 focus implies long sales cycles, small logo count early.
  - Average dataset size per order (images): enterprise ~100k–500k images/order, self-serve ~10k–50k; rationale: DMS/OMS automotive use cases require large labeled datasets.
  - Price per 1,000 images: $20–$80 depending on modality complexity (Faces vs. HIC vs. Objects); rationale: market benchmarks for synthetic data range $0.01–$0.10/image.
  - Platform / subscription fee: $0–$2k/month per seat (self-serve); enterprise via annual contract.
  - Gross margin: 65–75% at scale; rationale: cloud compute costs for simulation are material but should leverage at scale; comparable synthetic data cos. guide 70%+ at maturity.
  - S&M % of revenue: 40–50% in Year 1–2, declining to 25–30% by Year 4; rationale: enterprise direct sales heavy.
  - R&D % of revenue: 30–40% early (simulation R&D is core); rationale: 85+ engineers already on payroll.
  - Customer churn: net revenue retention 110–120% (expansion-led); rationale: ML teams increase dataset volumes as models improve.
  - Generator mix: Faces initially dominant → HIC ramps Year 2–3 → Objects in Context adds Year 3+.
  - Verticals: Automotive (DMS/OMS) and AR/VR as early revenue verticals.

- **Scenarios (Base / Bull / Bear - which variables flex):**
  - **Base:** 30% YoY new customer growth, 110% NRR, 70% gross margin by Year 3.
  - **Bull:** faster enterprise adoption (50% YoY), Objects generator launches Year 2, NRR 125%.
  - **Bear:** long enterprise sales cycles compress new logos (15% YoY), Faces-only revenue, margin pressure from GPU cost inflation.

- **Required sheets / outputs:**
  1. Assumptions - all drivers with / tags and scenario toggles.
  2. Revenue Build - by generator (Faces, HIC, Objects), by customer tier (enterprise / self-serve), usage volume × price.
  3. P&L (Income Statement) - revenue, COGS (cloud compute), gross profit, S&M, R&D, G&A, EBITDA.
  4. Headcount Plan - engineering/simulation, sales, G&A (feeds into R&D and S&M).
  5. Cash Flow & Runway - operating CF, capex (GPU infrastructure), cash balance.
  6. Customer Cohort - logo count, ACV, NRR waterfall by cohort year.
  7. Scenario Summary - Base / Bull / Bear on one tab.

## Frequently asked questions

### Is the Datagen financial model free?

Yes. The Datagen model is a free Excel download with live formulas.

### Can I change the assumptions?

Yes. The workbook is editable and its live formulas recalculate when assumptions change.
