Cassandra is famous for its raw write speed. Because it appends data sequentially, a properly tuned Cassandra cluster can ingest hundreds of thousands of writes per second with sub-millisecond latency.

When write latency spikes, it is almost never a CPU problem. It is a queueing problem. When the queues fill up, Cassandra executes backpressure, dropping your data on the floor to prevent the JVM from crashing.

Situation

In 2023, the 36-node Cassandra 4.1 cluster was absorbing ~150k write QPS perfectly—until cluster-wide write latency suddenly spiked over 500ms.

This latency cascaded upstream. Client drivers hit their WriteTimeoutException limits, and the message queue consumers stalled. nodetool tpstats showed thousands of pending and dropped tasks queuing up in MutationStage (there’s no separate commitlog thread pool in tpstats — commitlog I/O pressure shows up as MutationStage backpressure, since mutations block there waiting on the commitlog write), and disk monitoring (iostat) showed severe write latency on the EBS volume hosting the data directories.

The Problem

To debug a write timeout, you must understand Cassandra’s write path architecture.

When a write hits a replica, the data bifurcates into two distinct locations simultaneously:

  1. The CommitLog: An append-only log. Under the default periodic sync mode (covered below), this is written to the OS page cache and fsynced to disk on a timer rather than per-write, so it doesn’t guarantee durability against power loss for every individual write — only commitlog_sync: batch does that.
  2. The Memtable: An in-memory skip-list data structure.

The write is acknowledged as successful to the coordinator only after the mutation is applied to both the commitlog and the memtable in memory — under periodic sync this doesn’t mean the commitlog write has been fsynced to disk yet, just that it’s been handed to both structures.

flowchart TD
    Coord[Coordinator] -->|Mutation Request| Rep(Replica Node)
    
    Rep --> Fork{Bifurcate Write}
    
    Fork --> Memtable[Memtable — In-Memory Skip-List]
    Fork --> CommitLog[CommitLog — Sequential Disk Append]
    
    Memtable -->|Threshold Reached| Flush[Memtable Flush Writers]
    Flush --> SSTable[SSTable Data.db — Disk Flush]
    
    CommitLog -->|Acknowledged| Ack[Return Success]
    Memtable -->|Acknowledged| Ack
    
    %% Bottlenecks
    SSTable -.->|Disk IO Saturated| Flush
    Flush -.->|Queue Full| Memtable
    Memtable -.->|Memory Full| Backpressure[MutationStage — Dropped Messages]

The configuration suffered from two catastrophic flaws: First, we co-located the commitlog_directory on the same EBS gp3 volume as the data_file_directories. The commitlog demands pure, sequential I/O. When background compactions ran on the SSTables, they generated massive random read/write I/O, completely stalling the commitlog’s ability to append data.

Second, our memtable_flush_writers were undersized. When memtables fill up (based on memtable_heap_space_in_mb and memtable_offheap_space_in_mb), Cassandra queues them to be flushed to disk as SSTables. Because the disk was saturated by compaction, the flush writers backed up. When the flush queue is full, Cassandra cannot accept new writes into the memtable, triggering MutationStage dropped messages.

In Practice

The playbook executes a permanent physical separation of the I/O paths.

The architecture moves the commitlog to a dedicated, physically isolated EBS volume with a provisioned 3,000 IOPS and 250 MB/s throughput. This guaranteed that compaction could run at 100% capacity on the data volume without ever stealing an IOP from the write path.

We then tuned cassandra.yaml to optimize the flush queues:

# Dedicate specific threads to flushing memtables based on core count
memtable_flush_writers: 4

# Move memtable allocations off the JVM heap to prevent G1GC pauses
memtable_allocation_type: offheap_objects

Finally, the pattern builds a real-time Prometheus alert on CommitLogPendingTasks and MemtableFlushPendingTasks. If the queue builds up beyond 500 tasks, we receive an alert before client write timeouts begin.

Where It Breaks

Design ChoiceTradeoffMitigation
commitlog_sync: batchBatch mode forces a synchronous disk fsync on every write before acknowledging the client. It provides maximum durability but absolutely destroys throughput.Use periodic mode (sync every 10,000ms) for high-throughput workloads where OS page cache buffering is acceptable.
High offheap allocationSetting off-heap memtable space too high steals memory from the Linux OS page cache, degrading read performance.Balance memtable_offheap_space_in_mb against your system RAM, leaving at least 30% for the OS.
Combined Storage VolumesPutting commitlogs and data on one NVMe drive seems fast, but high concurrent load will inevitably cause write stalls.Always isolate the commitlog directory to a dedicated volume.

Modern Cassandra Context: Cassandra 5.0 and Trie Memtables

Cassandra 5.0 introduces Trie Memtables (memtable: trie). By replacing the legacy skip-list with a concurrent prefix tree, Cassandra significantly increases write concurrency. Because the Trie is vastly more efficient at managing memory overhead, it nearly eliminates the garbage collection stalls previously associated with high-velocity memtable churn, allowing operators to run massive write loads on smaller EC2 instances without dropping mutations.

What to Do Next

  • Problem: Co-locating commitlogs with SSTables causes compactions to starve the sequential write path, backing up flush queues and dropping mutations.
  • Solution: Move the commitlog to a dedicated EBS volume and shift memtable allocation off-heap.
  • Proof: Physical I/O separation isolates the write path, ensuring p99 write latencies remain stable regardless of background compaction storms.
  • Action: Run nodetool tpstats and check MutationStage. If Dropped is greater than 0, your nodes are silently executing backpressure. Check your iostat utilization immediately.