# Inside Asana's Database Architecture: Sharding and Scaling | Asana

> A technical deep dive into how Asana shards petabyte-scale customer data across 150+ MySQL instances, manages caching, and adapts its database layer for AI workloads.

Source: https://asana.com/inside-asana/database-architecture-sharding-scaling

## Inside Asana's Database Architecture: Sharding, Data Flow, and Scaling Tradeoffs

At Asana, changes to work appear in real time across the product by default. That responsiveness increases database traffic and pushes our databases toward their scaling limits as usage grows. LunaDb, our GraphQL-like declarative query system, loads data for the app. We’ve written [previously](https://asana.com/resources/scaling-lunadb) about how session management, data loading, and distributed cache invalidation work at scale.

This post looks at the database layer beneath LunaDb. Across our RDS MySQL instances, we store petabyte-scale customer data, with distributed caching helping us manage database traffic. We’ll explain how we model and shard that data, the tools we use to operate the databases at scale, and the tradeoffs involved.

### **The Asana Data Model**

To understand this, first, we need to understand how Asana’s data is laid out and sharded. Asana uses an [entity-attribute-value (EAV)](https://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80%93value_model) model to store customer data.

For example, if you have an object type Task with two fields, name: str and is_completed: bool, it will be stored like this.

We denormalize this data into physical index tables to make it easier for LunaDb to determine which cached data to invalidate. The mechanism by which this affects object &amp; query caching is out of scope for this discussion, but the important part is how it affects write fanout: a write to a single object value can cascade to many physical indices. 

### **How we distribute customer data across MySQL instances**

Customer data is sharded across well over a hundred physical MySQL instances, with the customer ID as the partition key. A customer’s data exists on a single MySQL instance¹, and databases are multi-tenant (i.e. many customers are mapped to one DB).

We intentionally have many customers share an instance in order to make our systems easier to use. One db per customer wouldn’t be practical at Asana’s scale.

If you want to read more about sharding, there is a blog post from the time Asana did that migration [here](https://asana.com/inside-asana/sharding-is-bitter-medicine).

Some data doesn’t neatly fit into the domain model. For example, an Asana account can be connected to foo.com and bar.com at the same time. How do you represent this object? We store this in a small set of databases that hold cross-domain data. Because these are shared widely, we’ve invested in reducing how much the rest of the system depends on them, including denormalizing data wherever possible.

Asana supports multiple data regions as well as a FedRAMP partition; our sharding strategy is the same across regions.

### **How we operate the database layer at scale**

Below are the major architectural decisions we made to operate our database layer reliably and efficiently at scale:

**We rely heavily on caching**_:_Asana has two different distributed caches (LunaDb and LunaServer Redis) that allow reduction of the DB layer complexity and reduce the need for massive instances or complexity around operating at the limits of MySQL’s transaction engine’s scale.

**We trade off some simple query performance in exchange for expressiveness**_:_Some of the simpler read operations that would otherwise just be a single row require gathering data from multiple tables eg. fields of a Task aren't in one place. This is a tradeoff we make to create more expressible queries with our graph query language.

**A transaction killer stops badly-behaved queries**_:_A long transaction killer is standard practice for operating an OLTP database. There are cases where Asana’s graph queries produce longer query times, and since the database has to keep track of extra copies of objects/connection overhead, a bunch of backed-up transactions can cause database performance to degrade rapidly. Asana's ORM allows users to manually open transactions; if app logic doesn't clean up, this can also cause DB overhead. 

**We backpressure async workloads**_:_ Asana’s biggest source of variable read/write load is from async work. We have a system, Infrastructure Resource Management, that tracks consumption of resources consumed by these jobs and applies throttling for these workloads to preserve the health of the database as well as other infra resources.

**We don’t use DB replicas (mostly)**_:_ It’s surprisingly complex to use replicas when implementing transactions with both reads and writes that safely interact with LunaDb, which has an eventual consistency model. There’s also a base level of operational and cost overhead that we’d pay for keeping this replica online – with how good LunaDb is, this layer is mostly not needed from a scaling perspective, so we’ve elected to avoid the additional complexity and maintenance cost of replicas². 

**We re-use connections via connection pooling**_:_ Asana has a mix of long-lived and ephemeral connections to each database. For the databases containing cross-domain data, there is a much larger number of ephemeral connections created and torn down. We use RDS Proxy as a connection pool to reduce the number of connections churning, which improved database stability &amp; also resulted in significant CPU savings.

**Asana rarely performs shard migrations within a region:**modern databases can scale vertically to a very large extent before they reach an actual scaling limit. However, the shard migration system is also used to move customers between Asana’s multi-region offerings, such as from the US to EU.

### **Where the architecture still gets strained**

**Spiky workloads**_:_ Spikes in uncached read traffic, scheduled asynchronous workloads, or unexpected query patterns can temporarily overwhelm databases. Our backpressure mechanisms and circuit breakers absorb most of this, and we keep tuning them.

**Operational Pain**_:_Low-downtime database engine upgrades are time-consuming right now, because LunaDb is heavily tied to physical binlog files and offsets.

**Hotspot shards**_:_ We have some larger enterprise customers that cause non uniform load distribution. An unusually large customer shard can change traffic patterns in a way that breaks existing scale assumptions, and we often handle these on a case-by-case basis.

### **How AI workloads are changing our scaling priorities**

With the AI era, there are a few fundamental shifts that we’ve seen.

AI agents are amplifying the data volume of reads and writes. This is happening for both throughput (agents work 24/7), and data storage (incremental cost to host). 

In the past, we’ve also been able to assume that compute, memory, and storage get _cheaper_ over time. That seems to no longer be true, and so another key aspect the team’s looking at is balancing new feature costs vs the hosting costs for the corresponding infrastructure. 

Addressing these challenges involves navigating complex engineering tradeoffs with no straightforward answers.

We also have some classical problems that the team is thinking about, such as how to preemptively “shift left” to catch query performance issues earlier in development, improving observability and attribution at the product feature level, and reducing toil for database engine changes. 

As we continue to scale Asana's infrastructure into the AI era, our focus remains on building resilient systems that balance performance, cost, and operational simplicity. By anticipating new demands and continuously refining our tooling, we ensure our database layer remains a solid foundation for every team and customer we support.

**Footnotes**

¹ Except for some metadata that doesn’t neatly fit the domain-sharded model

² For certain databases with extremely large customers, we do operate read replicas.

_Spencer Yu is the technical lead of the Core Storage Infrastructure team, which builds and manages the foundational cloud storage layer for all of Asana’s user data. This work was a group effort from the team and its past members, including but not limited to: Walter Li, Cynthia Gao, Debbie Pao, Claire Chen, David Hecker, Gaurav Ranade, Shreyas Patil, Shoaib Akbar, Ed Korthof, Rohan Batra, Sanchit Sinha, Mary Mathews, Bryant Lin, Jana Bantupalli_

- [Microframeworks in the Admin Console](/inside-asana/microframeworks-admin-console)

Engineering

Every Asana deployment has an admin console. It's where IT admins configure how their company uses Asana, such as adjusting password requirements, roles and permissions, whether f ...

- [Spec-driven development: The Good Parts - and what we learned after three months](/inside-asana/spec-driven-development)

Engineering

#### Staff Software Engineer

After three months, we had a clearer picture of when the added structure helped and when it got in the way.One of our engineers was preparing a data migration and decided to use s ...

- [We migrated off Enzyme in 2 weeks. It should have taken five years.](/inside-asana/migrating-off-enzyme-2-weeks)

Engineering

We recently used AI to complete years of engineering work in about one sprint. Here's how, and why it's changed how we think about what's possible.The five-year problemBack in 202 ...

- [Breaking the Lethal Trifecta: How Asana Thinks About Agentic AI Security](/inside-asana/how-asana-thinks-about-agentic-ai-security)

Engineering

#### Staff Security Engineer

Agentic AI introduces a class of security risk the industry hasn't solved. Here's how we think about it at Asana, and the security invariants we hold across our AI features.The pr ...

- [Inside Asana's Database Architecture: Sharding, Data Flow, and Scaling Tradeoffs](/inside-asana/database-architecture-sharding-scaling)

Engineering

At Asana, changes to work appear in real time across the product by default. That responsiveness increases database traffic and pushes our databases toward their scaling limits as ...

- [Engineering](/inside-asana/engineering-spotlight)
