How To Install Apache Cassandra on Ubuntu 26.04 LTS

Install Apache Cassandra on Ubuntu 26.04

A fresh Ubuntu 26.04 server, a quick apt install cassandra, and then nothing works. That is a common first experience on this release. The cause is almost never Cassandra itself. Ubuntu 26.04 LTS “Resolute Raccoon” made OpenJDK 25 the default Java version when you install default-jdk or default-jre. Cassandra 5.0 only supports Java 11 and 17. If you install packages without thinking about Java, you get a mismatch.

Installing Cassandra is easy. Installing it so it still behaves after a reboot, a traffic spike or a rolling upgrade takes more care. Many tutorials stop at “service is running” and call it done. Then the node falls over when the heap is wrong, swap is on, the firewall is open to the whole internet, or the cluster name was baked into the system keyspace before anyone edited the config.

This guide covers a clean Apache Cassandra installation on Ubuntu 26.04 LTS using the official Apache APT repository. It also covers the Java decision, first-start configuration, kernel and disk tuning, security, and the errors that tend to show up. The latest 5.0 release at the time of writing is 5.0.9, published on 2026-08-07. Everything below was written with that series in mind.

What You Need Before Starting

Be realistic about hardware. A laptop-sized VM works for learning. Production needs more.

Environment CPU RAM Disk Notes
Lab / learning 2 vCPU 4 GB 20 GB Reduce heap manually
Staging 4 vCPU 8 to 16 GB SSD/NVMe, 100 GB+ Mirror prod settings
Production node 8+ cores 32 to 64 GB NVMe, XFS or ext4 Separate commit log if on spinning disks

Other prerequisites:

  • Ubuntu 26.04 LTS with root or sudo access.
  • A static IP address or a DHCP reservation. Cassandra nodes that change IPs cause gossip trouble.
  • Working time synchronization. Cassandra uses timestamps to resolve write conflicts, so clock drift quietly corrupts your “last write wins” logic.
  • Outbound HTTPS to debian.cassandra.apache.org and downloads.apache.org.

Start with a full update:

sudo apt update && sudo apt full-upgrade -y
sudo reboot

Rebooting here is deliberate. If a kernel update is pending, it is better to land on it now than in the middle of a cluster bootstrap.

Choosing the Right Java Version

Java is the most common failure point on a new Ubuntu LTS. Cassandra 5.0 runs on Java 11 or Java 17. A build made with Java 17 cannot run on Java 11. Ubuntu’s default is 25, and the Cassandra 5.0 packages were not built for it. So install OpenJDK 17 explicitly.

sudo apt install -y openjdk-17-jre-headless ca-certificates curl gnupg
java -version

The headless JRE is enough for a server. You do not need the JDK, a GUI toolkit or the accessibility libraries that the full package drags in.

If java -version reports 25, or anything other than 17, another JDK is winning the alternatives race. This happens when something else on the box pulled in default-jre. Fix it like this:

sudo update-alternatives --config java
ls /usr/lib/jvm

Pick the entry for java-17-openjdk-amd64 (or arm64 on Graviton or Ampere hardware). Check again with java -version before going further.

Is Java 11 an option? Technically yes. But Java 17 gets longer support, has better G1 behavior, and is the version the project documents for 5.0. Pick 17 unless a compliance rule forces you otherwise.

Adding the Official Apache Cassandra Repository

Ubuntu’s own archive may carry a Cassandra package. Skip it. The Apache repository is the one that tracks upstream releases. The project documents this repository and key flow on its download page.

Import the signing key

sudo install -d -m 0755 /etc/apt/keyrings
sudo curl -fsSL -o /etc/apt/keyrings/apache-cassandra.asc https://downloads.apache.org/cassandra/KEYS

The signed-by approach keeps the key scoped to this one repository. The old apt-key add method is deprecated, and with good reason. A globally trusted key can sign packages for any repository on the system.

Add the repository

The suite name for the 5.0 series is 50x. For 4.1 it is 41x.

echo "deb [signed-by=/etc/apt/keyrings/apache-cassandra.asc] https://debian.cassandra.apache.org 50x main" | sudo tee /etc/apt/sources.list.d/cassandra.sources.list
sudo apt update

Note that the Debian and RedHat repository URLs moved, so any old sources.list entries from older blog posts need to be replaced. If you copied a config from an older server, check it for the previous hostname.

Before installing, check what APT is about to pull:

apt-cache policy cassandra

You should see a candidate in the 5.0.x range coming from debian.cassandra.apache.org. If the candidate is empty, the repository line has a typo, or the key failed to download.

Installing Apache Cassandra Without Starting It Blindly

Here is a trap. The Debian package starts the service immediately after install. It creates the data directories and writes the cluster name Test Cluster into the system keyspace on first boot. If you meant to rename the cluster, that is now a problem. Changing cluster_name later requires wiping system tables on a fresh node or following a more involved procedure.

On a node that will join a real cluster, block the auto-start first:

printf '#!/bin/sh\nexit 101\n' | sudo tee /usr/sbin/policy-rc.d
sudo chmod +x /usr/sbin/policy-rc.d

Then install:

sudo apt install -y cassandra

Remove the policy file once you have finished configuring:

sudo rm /usr/sbin/policy-rc.d

If you already let it start, stop the service and clear the data. Only do this on a node that holds nothing you care about:

sudo systemctl stop cassandra
sudo rm -rf /var/lib/cassandra/data/* /var/lib/cassandra/commitlog/* /var/lib/cassandra/saved_caches/* /var/lib/cassandra/hints/*

Pin the version afterward so a routine apt upgrade does not restart a database node unexpectedly:

sudo apt-mark hold cassandra

Rolling upgrades should be planned, one node at a time, with nodetool drain first. Treat them like a maintenance event and never let them happen by accident.

Where Everything Lives

The Debian package uses a predictable layout, which is handy when you automate with Ansible or Terraform.

Path Purpose
/etc/cassandra/ Configuration (cassandra.yaml, JVM options, snitch files)
/var/lib/cassandra/ Data, commit log, hints, saved caches
/var/log/cassandra/ system.log, debug.log, GC logs
/usr/share/cassandra/ Binaries and libraries
/etc/security/limits.d/cassandra.conf Resource limits for the cassandra user

Knowing these paths saves time at 3 a.m. when a disk fills up.

Configuring cassandra.yaml for a First Real Node

Open the main config:

sudo nano /etc/cassandra/cassandra.yaml

The defaults bind to localhost, which is fine for a single dev box and useless for a cluster. These are the settings that matter first.

Cluster identity

cluster_name: 'Prod-Cluster-ID1'
num_tokens: 16

Every node in a cluster must share the same cluster_name. Pick it once and keep it. The 16 token default in modern releases gives a better balance than the old 256 and makes repairs lighter.

Network addresses

listen_address: 10.10.1.11
rpc_address: 10.10.1.11

listen_address is used for node-to-node traffic. rpc_address is where clients connect. Never set either to 0.0.0.0 unless you also set broadcast_address or broadcast_rpc_address. Cassandra will refuse to start otherwise, and that refusal is correct behavior.

Seed nodes

seed_provider:
  - class_name: org.apache.cassandra.locator.SimpleSeedProvider
    parameters:
      - seeds: "10.10.1.11,10.10.1.12"

Two or three seeds per datacenter is enough. A node should not list only itself as a seed when it is meant to join an existing cluster. Otherwise it will happily bootstrap as a brand new, separate cluster. Many multi-node setups end up split this way.

Snitch

endpoint_snitch: GossipingPropertyFileSnitch

Use this one for anything beyond a toy setup. Then edit /etc/cassandra/cassandra-rackdc.properties:

dc=dc1
rack=rack1

Set the snitch before the node first joins a ring. Changing datacenter names later on a node with data triggers an error at startup and usually means a rebuild.

Authentication

Default Cassandra has no authentication. For anything touching a network:

authenticator: PasswordAuthenticator
authorizer: CassandraAuthorizer

Restart, log in with the default cassandra / cassandra account, create your own superuser, and then disable the default. More on this in the security section.

Heap and JVM Settings

Open the JVM options file that matches your Java version:

ls /etc/cassandra/jvm*-server.options
sudo nano /etc/cassandra/jvm-server.options

The common settings sit in jvm-server.options, and Java 17 specifics (G1 flags and so on) are in jvm17-server.options. For heap, uncomment and set matching minimum and maximum values:

-Xms8G
-Xmx8G

Why equal values? A dynamically resizing heap causes pauses at the worst time. Fix it at startup and let the JVM stay predictable.

A decent starting rule on G1 is roughly one quarter of system RAM, capped around 16 GB for most workloads. The rest of RAM is not wasted. Linux uses it for the page cache, which Cassandra leans on heavily for reads. Giving the heap 90 percent of memory is a classic mistake. It kills the page cache and the box gets slower.

A 4 GB lab VM can run with a 1 to 2 GB heap, which is enough to try out CQL.

Kernel and OS Tuning

This is the part that separates a demo from a production database.

Disable swap

Swap and Cassandra do not mix. A swapped-out JVM can stall for seconds, and gossip will mark the node as down.

sudo swapoff -a
sudo sed -i '/ swap / s/^/#/' /etc/fstab

If you must keep swap as a safety net, set vm.swappiness=1 instead, but outright removal is cleaner. Also check for a swap file with swapon --show, since Ubuntu cloud images often use one.

Kernel parameters

Create a sysctl file:

sudo tee /etc/sysctl.d/99-cassandra.conf <<'EOF'
vm.max_map_count = 1048575
vm.swappiness = 1
net.ipv4.tcp_keepalive_time = 60
net.ipv4.tcp_keepalive_probes = 3
net.ipv4.tcp_keepalive_intvl = 10
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
net.core.somaxconn = 4096
EOF
sudo sysctl --system

vm.max_map_count matters because Cassandra memory-maps SSTables. With the default value on a large dataset, the node can crash with out-of-memory errors even though plenty of RAM is free. The TCP keepalive values help detect dead connections faster between nodes.

File descriptor and process limits

The package ships a limits file, but confirm it:

cat /etc/security/limits.d/cassandra.conf

You want values along these lines:

cassandra - memlock unlimited
cassandra - nofile 1048576
cassandra - nproc 32768
cassandra - as unlimited

After the service runs, verify what the live process received:

cat /proc/$(pgrep -f CassandraDaemon)/limits | grep -E 'open files|processes|locked'

Do not assume. Systemd-managed services sometimes ignore limits.d, which is why checking the running process is the only trustworthy test.

Transparent Huge Pages

THP can cause latency spikes under memory pressure. Disable it at runtime and persist it with a small systemd unit or a kernel boot parameter (transparent_hugepage=never):

echo never | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/defrag

Disk and filesystem

XFS and ext4 both work well. Mount the data volume with noatime:

/dev/nvme1n1  /var/lib/cassandra  xfs  defaults,noatime  0 2

For SSD or NVMe, check the readahead and scheduler:

cat /sys/block/nvme1n1/queue/scheduler
sudo blockdev --getra /dev/nvme1n1

NVMe devices usually run with none already. Lower readahead (8 to 16 KB equivalent) suits random reads. On rotational disks, put the commit log on a separate spindle, because it is written sequentially and gets interrupted by SSTable flushes otherwise.

Time synchronization

timedatectl status
chronyc tracking

You want “System clock synchronized: yes” and offsets in milliseconds. Cluster nodes drifting by a second or more will produce odd read results that are extremely hard to diagnose after the fact.

Starting Cassandra and Checking Health

Now start the service:

sudo systemctl enable --now cassandra
sudo systemctl status cassandra --no-pager

Be patient. First start can take 30 to 90 seconds before the native transport port opens. Follow the log while waiting:

sudo tail -f /var/log/cassandra/system.log

You are waiting for a line saying the node is listening for CQL clients. Then run:

nodetool status

A healthy single node shows UN (Up, Normal) with an address, load, tokens and rack. If it shows DN, or the command hangs, go to the troubleshooting section.

Confirm the listening ports:

sudo ss -tlnp | grep -E '7000|7199|9042'
Port Purpose
7000 Inter-node communication
7001 Inter-node over TLS
7199 JMX (used by nodetool)
9042 CQL native transport (clients)

Connecting with cqlsh and Running a Real Test

cqlsh 10.10.1.11

If authentication is on, use cqlsh -u cassandra -p cassandra 10.10.1.11. Then try this in the shell:

CREATE KEYSPACE shop
  WITH replication = {'class': 'NetworkTopologyStrategy', 'dc1': 3};

USE shop;

CREATE TABLE orders (
  customer_id uuid,
  order_time timestamp,
  order_id uuid,
  total decimal,
  PRIMARY KEY ((customer_id), order_time)
) WITH CLUSTERING ORDER BY (order_time DESC);

INSERT INTO orders (customer_id, order_time, order_id, total)
VALUES (uuid(), toTimestamp(now()), uuid(), 149.90);

SELECT * FROM orders LIMIT 5;

Always use NetworkTopologyStrategy, even for a single datacenter. SimpleStrategy ignores topology, and migrating later is a headache. On a single test node, set the replication factor to 1, otherwise the keyspace will complain about unavailable replicas.

Cassandra 5.0 added features worth knowing about, including Storage-Attached Indexes (SAI) and a vector data type. A quick taste:

CREATE INDEX orders_total_idx ON shop.orders (total) USING 'sai';

SAI removes much of the old pain around secondary indexes, but data modeling around partition keys still comes first. Query-first design remains the rule. Define your read patterns, then build tables around them.

Securing the Node

An open Cassandra port is a common cause of data exposure. Treat security as part of the install, not a later task.

Firewall with UFW

Allow SSH first so you do not lock yourself out:

sudo ufw allow OpenSSH
sudo ufw allow from 10.10.1.0/24 to any port 7000 proto tcp
sudo ufw allow from 10.10.1.0/24 to any port 7001 proto tcp
sudo ufw allow from 10.10.2.0/24 to any port 9042 proto tcp
sudo ufw enable
sudo ufw status numbered

In this example, 10.10.1.0/24 is the database subnet and 10.10.2.0/24 is the application subnet. JMX on 7199 stays bound to localhost by default, and it should stay that way. Never expose it publicly.

Replace the default superuser

CREATE ROLE dbadmin WITH PASSWORD = 'use-a-long-random-passphrase' AND SUPERUSER = true AND LOGIN = true;

Log in as dbadmin, then:

ALTER ROLE cassandra WITH PASSWORD = 'another-long-random-value' AND SUPERUSER = false;

Then raise the replication of the system_auth keyspace so logins survive a node failure:

ALTER KEYSPACE system_auth
  WITH replication = {'class': 'NetworkTopologyStrategy', 'dc1': 3};

Run nodetool repair system_auth on each node afterward. Skipping this step is how clusters end up with nobody able to log in during an outage.

Encryption

Enable client-to-node and node-to-node TLS in cassandra.yaml under client_encryption_options and server_encryption_options. Generate certificates with your internal CA rather than self-signed one-offs, so rotation is manageable. Even inside a private VLAN, encryption is cheap insurance.

Permissions and updates

sudo chown -R cassandra:cassandra /var/lib/cassandra /var/log/cassandra
sudo chmod 750 /var/lib/cassandra

Run Cassandra as the cassandra user only. Keep the OS updated with unattended-upgrades for security patches, but exclude the held Cassandra package so database upgrades remain deliberate.

Multi-Node Cluster Notes

Adding a second and third node follows the same recipe: install Java 17, add the repository, install the package without auto-start, then apply the same cluster_name, seeds and snitch with that node’s own listen_address and rpc_address.

Bring nodes up one at a time. Wait until nodetool status shows the previous node as UN before starting the next. Starting several bootstraps at once can fail with token range conflicts. Patience here is cheaper than a rebuild.

After the cluster is stable, run nodetool cleanup on existing nodes if you added capacity, because they keep data for ranges they no longer own until told otherwise.

Troubleshooting Common Problems

cqlsh: Connection refused on 9042

Cassandra is probably still starting, or it bound to a different address. Check:

sudo ss -tlnp | grep 9042
grep -E '^(listen_address|rpc_address)' /etc/cassandra/cassandra.yaml

If the service listens on 10.10.1.11 and you run cqlsh localhost, it will fail. Connect using the configured address.

Service starts, then dies within seconds

Read the log, not the systemd status:

sudo journalctl -u cassandra -n 100 --no-pager
sudo tail -n 200 /var/log/cassandra/system.log

The usual causes are a wrong Java version, a heap larger than available RAM, or a syntax error in cassandra.yaml. YAML is picky about indentation, and tabs are not allowed.

Wrong Java version errors

Messages about unsupported class file versions or an unrecognized JVM mean Cassandra picked up Java 25 or Java 11 where it needs 17. Revisit update-alternatives --config java and confirm the version, then restart.

NO_PUBKEY or GPG errors during apt update

The key did not download or the path in signed-by is wrong:

ls -l /etc/apt/keyrings/apache-cassandra.asc

Re-download it and check that the repository line points at the same file. An empty file from a failed curl is a frequent culprit, which is why -f in the command matters.

Cannot start node if snitch’s data center differs from previous

You changed dc= after the node first started. Either revert the name or, on an empty node, wipe the data directories and start fresh. For a node with data, follow a proper rebuild.

Unable to gossip with any peers

Seeds are unreachable. Test the network and ports from the failing node:

nc -zv 10.10.1.11 7000

Look at UFW rules, cloud security groups and routes. Also verify that cluster_name matches exactly on all nodes, since a mismatch is rejected silently in some cases.

Too many open files

The limits did not apply to the running process. Check /proc/<pid>/limits as shown earlier. If they are low, set LimitNOFILE through a systemd override:

sudo systemctl edit cassandra
[Service]
LimitNOFILE=1048576
LimitMEMLOCK=infinity

Then sudo systemctl daemon-reload && sudo systemctl restart cassandra.

Long GC pauses and OOM kills

Check dmesg -T | grep -i kill for the kernel OOM killer. If the heap is set too high relative to RAM, the kernel may kill the JVM. Reduce the heap and remember the page cache needs room. Also look for oversized partitions with nodetool tablehistograms, which are often the real cause of memory pressure.

cqlsh fails with Python errors

cqlsh is a Python tool, and new Python releases sometimes break older bundled drivers. Check python3 --version, then read the release notes of your Cassandra version for supported Python versions. If you hit a wall, you can run cqlsh from another host or container with a compatible Python while the server remains untouched. The server itself does not depend on system Python.

r00t is an experienced Linux enthusiast and technical writer with a passion for open-source software. With years of hands-on experience in various Linux distributions, r00t has developed a deep understanding of the Linux ecosystem and its powerful tools. He holds certifications in SCE and has contributed to several open-source projects. r00t is dedicated to sharing her knowledge and expertise through well-researched and informative articles, helping others navigate the world of Linux with confidence.

Related Posts