What is Databricks?
Databricks is a data and AI company that provides a cloud-based platform for data engineering, machine learning, and analytics.
Company timeline
- Dec 2025 – Arvind Jain said initial attempts to automate software engineering at Databricks failed due to human organizational issues, not AI shortcomings.
- May 2026 – Jason Lemkin stated that at Databricks’ Neon (its Supabase competitor), over 90% of databases are built by agents, not humans.
- Jun 2026 – Reynold Xin said Databricks runs 50–60 million virtual machines a day across three clouds and processes exabytes of data before breakfast.
- Jun 2026 – Reynold Xin reported 13 million databases a day are created via an agentic factory built from a decade of traces.
- Jun 2026 – Matei Zaharia said Databricks has pipelines using open-source models that generate training environments and self-train to beat Opus and GPT 5.5 at tasks.
- Jun 2026 – Jason Lemkin and Rory O’Driscoll stated Databricks claims it can perform an enterprise-grade LLM migration in 30 days or less.
- Jul 2026 – Jason Calacanis said Databricks is among the major AI distribution players alongside SpaceX, OpenAI, and Anthropic.
- Aug 2026 – Nathaniel Whittemore noted that in a task comparison, Sonnet cost about $2.09 per task versus $1.94 for Opus, but because Sonnet required more iterations and tokens, the more expensive Opus was cheaper to operate overall.
Where it appears in the record
Every line below is attributed to a named speaker.
Why a cheaper-per-token model can cost more per task: lower-capability models require more iterations and reasoning tokens to reach the same result, making the nominally expensive model the economical choice at the task level.
“However, Sonnet cost around $2 per task or 2.09 per task versus 1.94 for Opus. So, because Sonnet needed more iterations and more reasoning had to spend way more tokens to get to the same results, overall Opus, which is significantly on paper more expensive model, it was cheaper to operate.”Nathaniel Whittemore · 4 Aug 2026
Databricks processes 50 to 60 million virtual machines per day across three clouds.
“We maybe 50 or 60 million virtual machines a day across our three clouds.”Reynold Xin · 24 Jun 2026
Over 90% of databases at Databricks Neon are created by AI agents, not humans.
“At data bricks neon, which is their superbase competitor, over 90% of the databases are built by agents, not by humans.”Jason Lemkin · 28 May 2026
Databricks Neon: a Supabase competitor where >90% of database creation is already agent-driven, making it a leading indicator of agentic infrastructure adoption.
“At data bricks neon, which is their superbase competitor, over 90% of the databases are built by agents, not by humans.”Jason Lemkin · 28 May 2026
Neon, acquired by Databricks, is scaling serverless Postgres to 13 million database launches per day, making it a key infrastructure bet on agent-era ephemeral compute needs.
“I think 13 million databases a day now.”Reynold Xin · 24 Jun 2026
Databricks argues that unifying storage (via L-TAP) delivers 99% of the benefits of HTAP without requiring a single database to handle both OLTP and analytics workloads simultaneously.
“We think you can get 99% of what you need by unifying the storage.”Reynold Xin · 24 Jun 2026
Opus 4.8 cost $1.94 per completed coding task vs $2.09 for Sonnet 5, despite Sonnet being cheaper per token (1.7x), because Sonnet required more iterations and reasoning tokens (Databricks benchmark).
“However, Sonnet cost around $2 per task or 2.09 per task versus 1.94 for Opus. So, because Sonnet needed more iterations and more reasoning had to spend way more tokens to get to the same results, overall Opus, which is significantly on paper more expensive model, it was cheaper to operate.”Nathaniel Whittemore · 4 Aug 2026
Neon (acquired by Databricks) launches approximately 13 million databases per day.
“I think 13 million databases a day now.”Reynold Xin · 24 Jun 2026
Arvind Jain says AI enterprise automation failures are a human organization problem, not an AI problem.
“Initial attempts at automate a lot of the software engineering at Databricks kind of failed even there's nothing wrong with the AI. The problem is the humans and how we were organized.”Arvind Jain · 23 Dec 2025
Open-source model self-training pipelines, where the same model generates its own training environments, can outperform frontier proprietary models like Opus and GPT-5.5 on specific tasks, challenging the assumption that frontier closed models hold a durable edge.
“We have pipelines just using open-source models. Like the same model generates training environments and trains itself and beats like Opus and GPT 5.5 and stuff at a task.”Matei Zaharia · 24 Jun 2026