Why We Moved from DataStax Enterprise to Apache Cassandra
In the early days of Cassandra, running the open-source version in production required a massive engineering team. DataStax Enterprise (DSE) solved this by bundling proprietary search, monitoring, and security plugins into a stable commercial offering.
However, as the Apache Cassandra project matured through the 4.x lifecycle, the open-source community built enterprise-grade tooling that rivaled and often surpassed proprietary alternatives. When the massive DSE licensing renewal arrived, we realized we were paying a premium for vendor lock-in.
Situation
By 2024, the 45-node DSE 6.8 cluster (~40 TB dataset) was the backbone of the platform. The commercial licensing and support subscriptions had become one of the highest infrastructure line items.
We were mandated to migrate to 100% open-source Apache Cassandra 4.1. This was not a simple binary swap. We were deeply coupled to DSE-specific features: DSE Search (Solr), DSE Graph, DSE Analytics (Spark), OpsCenter for monitoring, and the proprietary com.datastax.dse.driver in the Java microservices.
Attempting to boot Apache Cassandra directly on DSE SSTables would fail because DSE uses proprietary header identifiers and on-disk format modifications. Operators must execute a live, zero-downtime migration across 40 TB of data.
The Problem
Migrating off a proprietary database fork requires auditing and replacing every vendor-locked dependency before you can touch the data.
- Observability: Operators must replace OpsCenter, the proprietary GUI previously used for monitoring and repair orchestration.
- Search: Operators must decouple the application from DSE Search (Solr), which allowed developers to run
LIKEqueries and text searches directly against Cassandra tables. - The Data Migration: Because the SSTable formats were incompatible, we could not use
nodetool upgradesstables. Operators must physically stream 40 TB of data from the old cluster to the new cluster while the application was still serving live traffic.
The Open-Source Migration Blueprint
The pattern architects a 4-Phase Migration Blueprint that relied on dual-writing and asynchronous data backfilling.
flowchart TD
App[Application Microservices] --> Proxy{Dual-Write Proxy}
subgraph "Legacy DSE 6.8 Cluster"
Proxy -->|Primary Write — Read| DSE[DSE Nodes]
end
subgraph "New OSS Apache Cassandra 4.1 Cluster"
Proxy -.->|Shadow Write| OSS[OSS Nodes]
DSBulk[DSBulk Pipeline] -->|Historical Backfill| OSS
end
DSE -->|Source Data| DSBulk
%% Decoupled Search
Proxy -->|Search Queries| OS[OpenSearch Cluster]
1. Decoupling Search and Observability
First, we eliminated the proprietary DSE Search dependency. We stood up an Amazon OpenSearch cluster and updated the application to index search-heavy data directly into OpenSearch, removing the reliance on DSE Solr cores.
To replace OpsCenter, a typical deployment involves the open-source “Holy Trinity” of Cassandra operations: the Prometheus Cassandra Exporter for metrics, Grafana for dashboards, and Cassandra Reaper to automate anti-entropy repair scheduling.
DSE Search and OpsCenter were the tractable dependencies. DSE Graph and DSE Analytics (Spark) were not: neither has a drop-in OSS Cassandra equivalent. The graph workload was rebuilt as a separate service on a purpose-built graph database, reading from Cassandra via CDC rather than living inside it, and the Spark analytics jobs were repointed at the OSS cluster using the open-source Spark Cassandra Connector — both were multi-quarter efforts running in parallel with, not inside, the migration blueprint below. The proprietary com.datastax.dse.driver authentication and LDAP integration also had no equivalent in OSS Cassandra’s PasswordAuthenticator; that gap was closed by fronting the cluster with an external identity-aware proxy rather than relying on database-native LDAP binding.
2. The Dual-Write Abstraction
We provisioned a brand new, parallel Apache Cassandra 4.1 cluster. We updated the application to use the open-source driver (com.datastax.oss.driver) and implemented a dual-write proxy. The application wrote data synchronously to the legacy DSE cluster (primary) and asynchronously to the new OSS cluster (shadow). Reads were served exclusively from the DSE cluster.
3. Historical Backfill with DSBulk
While the application was dual-writing live traffic, the new OSS cluster was still missing historical data. The pipeline utilizes DataStax Bulk Loader (DSBulk), an open-source CQL-based data transfer tool, to extract the 40 TB of historical data from the DSE cluster and load it directly into the OSS cluster. Because DSBulk operates via CQL, it bypassed the incompatible binary SSTable formats completely.
In Practice
Once the DSBulk backfill completed, both clusters contained identical data.
To prove data integrity, we enabled a “Shadow Read Validator” in the application. When a read request came in, the application fetched the data from DSE and returned it to the client. Asynchronously, it fetched the exact same data from the OSS cluster and compared the payloads. We monitored the discrepancy metrics for two weeks.
Once the discrepancy rate hit absolute zero, the playbook executes a zero-downtime DNS cutover, routing all primary reads and writes to the open-source Apache Cassandra cluster. We kept the legacy DSE cluster running as a dual-write shadow for 30 days as a rollback safety net before finally decommissioning it.
Where It Breaks
| Design Choice | Tradeoff | Mitigation |
|---|---|---|
| Direct SSTable Copying | Moving DSE SSTables directly into Apache Cassandra will crash the node on startup due to unreadable file headers. | You must use CQL-based migration tools (like DSBulk) or dual-write replication to move data between forks. |
| Missing Driver Breaking Changes | The Java Driver 3.x to 4.x migration involves significant API changes for asynchronous queries and load balancing policies. | Allocate adequate sprint capacity to rewrite driver connection logic. It is not a drop-in JAR replacement. |
| Skipping Repair Before Cutover | If the new OSS cluster drops a mutation during the dual-write phase, the data will be silently inconsistent upon cutover. | Run a full Cassandra Reaper repair cycle on the new cluster before cutting over the primary reads. |
Modern Cassandra Context: Cassandra 5.0 Closes the Feature Gap
Operating pure open-source Apache Cassandra has drastically reduced the Total Cost of Ownership (TCO). Furthermore, the release of Apache Cassandra 5.0 closes the final feature gaps that historically drove teams to DSE. With the introduction of Storage-Attached Indexing (SAI), open-source Cassandra now supports vector search and high-performance secondary indexing natively, eliminating the need to bolt on external search engines for basic lookup patterns.
What to Do Next
- Problem: Vendor lock-in to proprietary Cassandra forks prevents infrastructure agility and incurs massive licensing costs.
- Solution: Deploy a parallel open-source cluster, implement application-level dual-writing, and use DSBulk to migrate historical data via CQL.
- Proof: Shadow-read validation ensures absolute data parity between the legacy and open-source clusters before executing a zero-downtime DNS cutover.
- Action: If you are migrating drivers, upgrade to
com.datastax.oss.driverimmediately. It natively supports both DSE and OSS clusters, allowing you to modernize your application code before moving the data.
Interactive tools for this topic