Linux and EC2 OS Tuning for Cassandra
A database is only as reliable as the operating system underneath it. You can spend weeks tuning Cassandra’s JVM and compaction strategies, but if the Linux kernel is fighting the database for memory pages, your cluster will randomly freeze.
Situation
In 2022, a typical deployment involves a 27-node Cassandra 4.0 cluster on AWS EC2 using a standard Ubuntu 20.04 LTS AMI. Within weeks, clusters often experience unexplained node freezes.
During peak traffic, nodes would lock up for 10–15 seconds. system.log was flooded with java.lang.OutOfMemoryError: Map failed errors, even though the JVM heap had plenty of free space. Disk I/O stalled, and OS monitoring showed CPU spikes entirely consumed by the kernel’s mm/vmscan process. Furthermore, swap memory usage was climbing even though the instances had 64 GB of physical RAM.
The Problem
Out of the box, standard Linux distributions are tuned for generic, multi-tenant workloads. They attempt to optimize memory for desktop applications or stateless web servers. Cassandra is a stateful database that bypasses much of this logic to manage memory and disk directly.
When default Linux settings collide with Cassandra’s architecture, three catastrophic things happen:
- Transparent Huge Pages (THP): Linux tries to merge 4KB memory pages into 2MB blocks to improve CPU cache hits. When memory fragments, the kernel locks the system to defragment it. Cassandra’s allocation patterns guarantee memory fragmentation, causing THP to freeze the OS.
- Memory Mapped Files (mmap): Cassandra maps its SSTables directly into memory to allow the OS page cache to handle reads. A default Linux kernel limits a process to 65,530 memory maps (
vm.max_map_count). When Cassandra opens thousands of SSTables, it hits this limit and crashes with a map failure. - OS Readahead: Linux assumes sequential disk reads and pre-fetches data. If
readaheadis set to 128KB, a Cassandra random 4KB read fetches 128KB of data from EBS, wasting 96% of the I/O throughput and saturating the EBS network link.
The Kernel Storage and Memory Architecture
To fix this, operators must systematically disable the OS “optimizations” and hand control back to Cassandra.
flowchart TD
Cass[Cassandra JVM] -->|mmap SSTable| PageCache[Linux Page Cache]
Cass -->|Direct IO| NVMe[EC2 NVMe — EBS Driver]
subgraph "Linux Kernel Memory Management"
PageCache --> THP{Transparent Huge Pages}
THP -->|Enabled: Defrag Lock| Freeze[OS Freeze — CPU Spike]
THP -->|Disabled: 4KB Pages| FastMem[Stable Allocations]
PageCache --> Swap{Swappiness > 0?}
Swap -->|Yes| DiskSwap[Pages written to Disk]
Swap -->|Swappiness 0 or 1| RAM[Pages kept in RAM]
end
1. Disable Transparent Huge Pages
You must disable THP completely. A common mitigation is to implement a systemd service that executes echo never > /sys/kernel/mm/transparent_hugepage/enabled before the Cassandra daemon starts. This instantly eliminated the mm/vmscan CPU spikes.
2. Tuning Sysctl for Databases
We created a hardened /etc/sysctl.d/99-cassandra.conf file to override default kernel behavior:
vm.swappiness = 1: This tells the kernel to aggressively keep data in RAM and only use swap to prevent an OOM-killer panic. (Do not set it to0, or the kernel will violently kill processes under memory pressure).vm.max_map_count = 1048575: This provides the headroom Cassandra needs to map thousands of SSTable components.
3. Disk Readahead
Apache Cassandra’s production baseline for random-read workloads is a 4 KB readahead. We enforced this at the block device level using blockdev --setra 8 /dev/nvme1n1 (8 sectors × 512 bytes = 4 KB). This prevented the OS from saturating the EBS bandwidth by pre-fetching data Cassandra would never use.
In Practice
The documented pattern for deploying Cassandra on Linux is to codify OS overrides into a configuration management tool like Ansible before the database daemon is ever installed.
Because of how the Linux page cache and TCP stack behave under high concurrency, operators apply a strict set of kernel tunings to override default generic-server behaviors:
# /etc/sysctl.d/99-cassandra.conf
vm.swappiness = 1
vm.max_map_count = 1048575
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216
To prevent the kernel from terminating the JVM under memory pressure, the documented pattern also configures PAM limits (/etc/security/limits.d/cassandra.conf) to allow the cassandra user to lock memory (memlock) and open up to 1048576 file descriptors (nofile).
Cassandra’s behavior when these limits are misconfigured is catastrophic but predictable. When max_map_count is too low, the database cannot map SSTable components into memory, resulting in OutOfMemoryError: Map failed exceptions that force the JVM to crash. Furthermore, without sufficient file descriptors, the database cannot maintain its connection pools or open required SSTables during compaction, leading to failed reads and artificial I/O pressure as the system retries failed memory maps. Applying these limits strictly, alongside blockdev --setra 8 for 4 KB disk readahead, stabilizes the allocation and storage paths.
Where It Breaks
| Design Choice | Tradeoff | Mitigation |
|---|---|---|
| Disabling Swap Completely | If you disable swap entirely (swapoff -a), a sudden memory spike will cause the kernel OOM-killer to instantly terminate the Cassandra process. | Keep swap enabled but set vm.swappiness=1 for graceful degradation under extreme pressure. |
| High OS Readahead | A 128KB readahead causes massive read amplification on EBS volumes for random 4KB data fetches. | Strictly enforce blockdev --setra 8. |
| Default CPU Scaling | Standard AMIs use the ondemand CPU governor, which introduces microsecond latencies when ramping up clock speeds. | Set the scaling governor to performance where the instance family actually exposes scaling_governor control to the guest OS — check /sys/devices/system/cpu/cpu0/cpufreq/ first, since this varies by instance generation and hypervisor. |
Modern Cassandra Context: Linux Kernel 6.x and NVMe Schedulers
Modern EC2 instances running Linux Kernel 6.x and AWS Nitro have changed the storage I/O stack. With the introduction of io_uring considerations and modern NVMe drivers, operators must ensure that the block layer I/O scheduler is set to none (bypassing legacy scheduling) rather than mq-deadline, allowing the Nitro hardware controller to manage queue depth natively. Furthermore, Cassandra 5.0’s startup scripts (cassandra-env.sh) now actively validate these sysctl baselines and will warn operators before the daemon starts.
What to Do Next
- Problem: Default Linux AMIs enable Transparent Huge Pages and restrictive memory map limits, causing memory fragmentation stalls and process crashes on stateful databases.
- Solution: Disable THP, set
vm.swappiness=1, increasevm.max_map_count, and clamp disk readahead to 4 KB. - Proof: Hardening these OS layers instantly stops
vmscanCPU spikes and eliminates EBS bandwidth saturation caused by useless pre-fetching. - Action: SSH into a node and run
cat /sys/kernel/mm/transparent_hugepage/enabled. If the output shows[always]instead of[never], your cluster is a ticking time bomb.