Atlassian’s Ambitious Migration: Jira Databases to Amazon Aurora
In a significant technological leap, Atlassian recently undertook the mammoth task of migrating 4 million Jira databases to Amazon Aurora. This strategic move aims to reduce costs and bolster the reliability of the Jira Cloud platform, making it an intriguing case study for enterprises facing similar challenges in cloud migration.
The Unique Architectural Approach
Atlassian’s architecture for Jira is noteworthy for its decision to allocate one database per tenant. While this model is generally considered impractical with larger tenant counts, it has vital advantages at scale. Pat Rubis, a principal site reliability engineer at Atlassian, details the rationale behind this choice:
“One database per tenant is an uncommon architecture… It makes it much easier to ensure that data from one tenant cannot accidentally or maliciously be accessed by another.”
This architecture allows for significant isolation, optimized load balancing, and enhanced operational control. However, it also introduces complexities; the team sometimes needs to rebalance the databases to maintain consistent performance across instances.
The Migration Process
As Atlassian prepared for a replatforming to Amazon Aurora, it aimed to capitalize on several benefits: a better SLA (99.99%), improved elasticity through the autoscaling of reader instances, and potential cost optimizations. The project was estimated to unfold over several months, with a particular focus on minimizing tenant downtime and migration costs.
Utilizing AWS Step Functions, Atlassian orchestrated the migration process, supported by feature flags to quickly update tenants’ database endpoints on application servers. Despite the straightforward nature of converting an Amazon RDS for PostgreSQL instance to Aurora, the large number of databases per instance presented unique challenges. Specifically, traditional migration methods wouldn’t work for the multitude of tenants with varying connection endpoints.
Overcoming Technical Challenges
Each Jira database corresponds to roughly 5000 files on disk, resulting in millions of files per PostgreSQL instance. This abundance posed a challenge, as Aurora’s limitations became apparent during the migration. New replica instances were timing out, hindering safe cluster conversions.
To tackle this, the team developed a method termed “draining.” This approach involved gradually reducing the number of tenants on instances slated for conversion, which controlled the number of databases being transferred across clusters, effectively mitigating the load and preventing timeouts.
Managing Concurrency
A primary concern throughout the migration process was the concurrency of both source and destination databases. Balancing the need for additional infrastructure in each region against the desire to keep costs low became a fine-tuned orchestration. Rubis noted:
“Ultimately, we had to find a balance between how much additional infrastructure we wanted… and how long we were comfortable with each region taking to complete.”
At peak capacity, Atlassian successfully migrated up to 90,000 Jira databases per day, averaging around 38,000 databases daily. The scale of this migration is significant, demonstrating Atlassian’s commitment to enhancing their infrastructure.
The Outcome and Implications
The migration encompassed a total of 2403 RDS database instances, facilitating the transfer of 2.6 million databases, with an additional 1.8 million databases drained from the original instances. The project culminated in the management of an estimated 27.4 billion database files within Jira, showcasing not just the scale of the operation, but also the technical prowess required to execute it.
Notably, Atlassian has chosen not to disclose specific metrics on cost savings realized through this migration, but the anticipated benefits in scalability, reliability, and efficiency are undeniably promising.
Key Takeaways
Atlassian’s migration story serves as an essential reference for organizations contemplating a large-scale cloud transition. Their innovative approaches to managing massive data clusters while addressing unique architectural challenges underscore the importance of planning, phase management, and technology adoption during such endeavors.
In the ever-evolving landscape of cloud solutions, Atlassian’s experience with Amazon Aurora exemplifies the potential for improved operational efficiencies through careful engineering and strategic foresight. Such insights can inspire other teams to explore similar pathways in their modernization efforts.
Inspired by: Source

