# Databricks Series D Financial Model

Databricks is the unified analytics platform (Serverless Spark PaaS + Data Science Workspace SaaS) that bridges the gap between big data infrastructure and AI/ML applications.

- Canonical: https://finamodel.com/startups/databricks-series-d
- Excel download: https://finamodel.com/startup-models/databricks-series-d.xlsx
- Category: AI/ML
- Model type: SaaS ARR / Valuation
- Funding round: Series D
- Funding: $140M
- Founded: 2017
- Geography: Global (cloud-agnostic: AWS, Azure, GCP) with on-prem expansion planned. [DECK slides 15, 17]
- Customer: B2B

## About the company

Databricks combines a serverless Spark platform with a collaborative Data Science Workspace. The cloud-agnostic product runs across AWS, Azure, Google Cloud, and on-premises environments, offering governed data connectors, notebooks, dashboards, and reports for enterprise analytics and AI workloads.

At the Series D deck date, Databricks had more than 500 customers, $21.9 million ARR, $118,000 average customer ARR, 136% net dollar retention, and 69% gross margin. Its Q1 2017 ARR was growing 248% year on year, with $36.7 million forecast for Q4.

The model uses quarterly ARR waterfalls for annual enterprise contracts: new customers, expansion, and churn. It translates ARR into recognised revenue and cloud COGS, models retention, sales efficiency and LTV:CAC, then forecasts technical headcount, operating cash flow, and deployment of its $56.1 million cash balance.

## What's included

- 5-year monthly revenue build with stage-appropriate growth assumptions
- Full P&L, headcount plan, and operating-expense schedule
- Cash-flow statement, runway, and burn-rate tracking
- Valuation via exit multiple with a DCF cross-check
- Returns analysis with MOIC and IRR
- Unit economics including CAC, LTV, payback, and cohort retention

## Product & value proposition

- Two-layer platform:
  - **Databricks Serverless Spark Platform** (PaaS): auto-tuning Spark with 10–40x speedup vs. open-source Apache Spark; custom data connectors; SOC2 & HIPAA compliant governance.
  - **Databricks Data Science Workspace** (SaaS): collaborative notebooks, real-time dashboards, periodic reports - democratizes Spark access for data teams.
- Cloud-agnostic: runs on AWS, Azure, GCP, and on-prem (YARN/Cloudera/Hortonworks/Kubernetes).
- Built by the original Apache Spark creators at UC Berkeley; Spark is the de facto standard for enterprise AI/ML at scale.
- Addresses the "AI gap": the hardest part of AI is big-data infrastructure (Google NIPS 2015 framing).

## Market

- Cloud computing market: Gartner estimate of $200B by 2020.
- Big data: "90% of the data created in last 2 years."
- ML/AI: described as having "just scratched the surface of use-cases."
- Spark ecosystem traction (proxy for addressable market): Summit attendees 1,100 (2014) → 3,900 (2015) → 5,100 (2016); Meetup members 12K (2014) → 66K (2015) → 300K+ (2016).
- No formal TAM/SAM/SOM breakdown provided.

## Revenue model

- Subscription ARR - enterprise contracts billed annually. Implied by ARR metrics and average customer ARR.
- Average Customer ARR = $118K as of 3/31/17.
- 500+ customers across all major verticals (Ad/Marketing, Media, Healthcare/Pharma, Enterprise Software, Public Sector, Financial Services, Industrial/IoT, Retail/CPG).
- Channel: direct enterprise sales (CRO background: Axway, Stanford MBA; CMO: Alteryx pre-IPO).
- Partnership/cloud marketplace channel mentioned via Michael Hoff (SVP BD, ex-Tableau, Azure).

## Traction & metrics

- ARR: $21.9M as of 3/31/17
- ARR annual growth: 248% (3/31/17 vs. 3/31/16)
- 5x growth Q1'17 vs. Q1'16
- ARR quarterly progression (actuals + forecast):
  - Q1'16: $6.3M | Q2'16: $7.9M | Q3'16: $12.3M | Q4'16: $16.7M
  - Q1'17: $21.9M | Q2'17: $25.6M (forecast) | Q3'17: $31.0M (forecast) | Q4'17: $36.7M (forecast)
- T2D3 comparison: T2D3 target ARR Q2'17 = $18M vs. Databricks actual/forecast = $25.6M - tracking ahead of T2D3.
  - Q2'15: T2D3 $2M / Databricks $1.9M; Q2'16: T2D3 $6M / Databricks $7.9M; Q2'17: T2D3 $18M / Databricks $25.6M
- Customer count: 500+
- Average Customer ARR: $118K (up 112% yr/yr)
- Net Dollar Retention: 136% (3/31/17 vs. 3/31/16)
- Cash on hand: $56.1M as of 3/31/17
- Gross Margin: 69% (Q1'17)
- LTV:CAC ratio: 4.1 (3/31/17 TTM)

## Unit economics

- Gross Margin: 69% Q1'17
- LTV:CAC: 4.1 (TTM as of 3/31/17)
- Net Dollar Retention: 136% - strong expansion motion; existing customers grow significantly.
- Average Customer ARR: $118K, up 112% yr/yr - driven by both new logos and expansion.
- Payback period: Not explicitly stated. With 69% GM and LTV:CAC of 4.1, implied payback ~14–18 months (typical for enterprise SaaS at this NRR), but not confirmable from deck.

## Competition / moat

- Moat framing: Databricks created Apache Spark (Berkeley origins); deepest technical expertise in the ecosystem.
- Platform lock-in: collaborative workspace + serverless infra = stickiness across data engineering, data science, and analytics teams.
- Cloud-agnostic positioning differentiates from single-cloud-native offerings (AWS EMR, Azure HDInsight, GCP Dataproc).
- On-prem expansion addresses Hadoop/Cloudera/Hortonworks incumbents.
- Named competitors: Not explicitly named in deck. Implicitly: legacy data warehouses (Netezza, Oracle), Hadoop vendors (Cloudera, Hortonworks, MapR), cloud storage.
- "AI gap" framing positions Databricks as the only unified solution across ETL, ML, and streaming.

## Team & funding ask / use of funds

- Team:
  - Ali Ghodsi - CEO & co-founder; PhD/MBA, Adjunct Professor UC Berkeley
  - Patrick Wendell - VP Engineering & co-founder; UC Berkeley MSc, Princeton BS
  - Ron Gabrisko - CRO; Cyclone (pre-revenue → Axway acquisition), Stanford MBA/MS
  - Rick Schultz - CMO; Alteryx (pre-revenue → IPO), Oracle VP Product Marketing
  - John Winkenbach - SVP Finance; CFO Jobvite, VP Finance Technorati
  - Hatim Shafique - CCO; AppDynamics (pre-revenue → IPO)
  - Michael Hoff - SVP BD & Partnerships; Tableau VP Channels, EMC Global VP Sales, MSFT GM Windows Azure
- Cash on hand pre-raise: $56.1M (3/31/17).

---

## Recommended financial model

- **Archetype + why:** SaaS ARR subscription model with usage-based expansion layer. Databricks sells annual enterprise contracts with strong net dollar retention (136%), so the correct model is an ARR waterfall (new ARR + expansion ARR − churn ARR) driving a P&L, not a transactional revenue model. The 69% gross margin, LTV:CAC of 4.1, and cloud infrastructure cost structure are all consistent with a cloud SaaS P&L.

- **Forecast horizon & granularity:** Quarterly for Years 1–2 (2017–2018), annual for Years 3–5 (2019–2021). The deck already provides quarterly ARR actuals through Q1'17 and management forecasts through Q4'17, so Q-level is natural.

- **Key drivers & assumptions:**

| Driver | Value | Source |
| -- | -- | -- |
| ARR at model start (Q1'17) | $21.9M | - |
| ARR annual growth rate (historical) | 248% | - |
| Q4'17 ARR target | $36.7M | - |
| Average Customer ARR | $118K | - |
| Customer count | 500+ | - |
| Net Dollar Retention | 136% | - |
| Gross Margin | 69% | - |
| LTV:CAC | 4.1 | - |
| ARR growth rate Year 2 (2018) | ~150% | Deceleration from 248% consistent with T2D3 double-double-double trajectory and scale |
| ARR growth rate Year 3 (2019) | ~100% | Continued scale; Series D capital deployed into S&M |
| ARR growth rate Years 4–5 | 60–80% | Normalization as base grows; comparable SaaS at $100M+ ARR |
| NDR long-term | 120–130% | Slight moderation from 136%; still strong expansion motion |
| Gross margin long-term | 72–75% | Modest improvement as platform scales; cloud infra costs optimized |
| S&M % of revenue | 45–55% | Typical high-growth enterprise SaaS; consistent with LTV:CAC of 4.1 |
| R&D % of revenue | 25–35% | Deep technical product; Spark ecosystem investment |
| G&A % of revenue | 8–12% | Standard enterprise SaaS overhead |
| New customer adds per quarter | ~30–50 | Implied from 500+ customers over ~10 quarters of operation |
| Average contract length | 1 year | Standard enterprise SaaS; not stated in deck |

- **Scenarios (Base / Bull / Bear - which variables flex):**
  - **Base:** ARR growth decelerates per table above; NDR holds at ~128%; GM improves to 73% by Year 3.
  - **Bull:** NDR sustains at 136%+; on-prem expansion adds a new revenue stream; cloud partnerships accelerate new logo adds; ARR growth stays >100% through 2019.
  - **Bear:** Enterprise sales cycle elongates; ARR growth decelerates to ~80% in 2018; NDR compresses to 115% as budget scrutiny increases; GM pressure from cloud infra costs.
  - Flex variables: new ARR growth rate, NDR, gross margin, S&M efficiency (CAC).

- **Required sheets / outputs:**
  1. **Assumptions** - all drivers in one place, color-coded vs.
  2. **ARR Waterfall** - beginning ARR, new ARR, expansion ARR, churn ARR, ending ARR (quarterly then annual)
  3. **Customer Cohort Schedule** - new logos per period, average ACV, expansion rate by cohort
  4. **P&L** - revenue (ARR → recognized revenue), COGS, gross profit, S&M, R&D, G&A, EBITDA, net income
  5. **Cash Flow** - operating cash flow; key: deferred revenue build (annual prepay), capex light
  6. **KPI Dashboard** - ARR, ARR growth %, NDR, gross margin, LTV:CAC, average customer ARR, cash runway
  7. **Scenario Toggle** - Base / Bull / Bear switcher driving all sheets

## Frequently asked questions

### Is the Databricks Series D financial model free?

Yes. The Databricks Series D model is a free Excel download with live formulas.

### Can I change the assumptions?

Yes. The workbook is editable and its live formulas recalculate when assumptions change.
