Kubernetes will prevent kubelet from starting on nodes with cgroup v1 starting in version 1.35
A post on the official Kubernetes blog, published on October 6, details the cgroup v1 deprecation timeline: starting with version 1.35, the kubelet simply won't come up on nodes still using the old model, unless someone forces a temporary override.
Deprecating a Kubernetes feature is rarely a single event: it's a process that drags on across several minor releases until it becomes a real blocker. That's the case with cgroup v1. A post on the official Kubernetes blog, signed by Paco Xu (DaoCloud) and published this past Tuesday (October 6), rounds up the current state of this transition, and the point that deserves attention from anyone running a cluster in production is concrete: starting with version 1.35, the kubelet won't start on a Linux node that still runs cgroup v1, unless someone explicitly configures an override.
The timeline already underway
cgroup v2 is nothing new: it has been stable in Kubernetes since version 1.25. What changed was the treatment given to v1. In version 1.31, cgroup v1 support entered maintenance mode, which in practice means the project stops investing in improvements and starts treating v1 as legacy. Starting with version 1.35 (already released), the failCgroupV1 flag now defaults to true, and this is where the breaking change actually shows up: unless that node is on cgroup v2, the kubelet fails to start.
Anyone operating via kubeadm feels the block even earlier. Starting with 1.35, the SystemVerification preflight check (from the k8s.io/system-validators package) starts returning an error, not just a warning, when it detects cgroup v1 combined with kubelet 1.35 or newer, during kubeadm init, kubeadm join, and kubeadm upgrade. With an older kubelet, the check remains just a warning. Complete removal of the fallback is scheduled for version 1.38 and is tracked by KEP-5573.
In short: if your cluster is currently running a version earlier than 1.35, migrating the Linux nodes to cgroup v2 needs to happen before the upgrade, or someone will have to set failCgroupV1: false in the kubelet configuration file as a temporary workaround, knowing that this workaround also has an expiration date.
What actually needs to be in place
Migrating to cgroup v2 isn't just a kernel flag. The post details the infrastructure requirements that need to be aligned before any upgrade:
| Component | Minimum requirement |
|---|---|
| Linux kernel | 5.8 or newer (5.9+ recommended if using Memory QoS) |
| Kubernetes | Any currently supported version (v2 stable since 1.25) |
| containerd | v1.4+ supports cgroup v2; v2.0+ brings automatic cgroup driver detection |
| CRI-O | v1.20 or newer |
| cAdvisor | v0.43.0 or newer, for those reading the cgroup filesystem directly |
Missing any row in this table won't bring the cluster down immediately, but it creates silent inconsistency: monitoring tools that read /sys/fs/cgroup directly, for instance, can simply stop finding the files they expected, because the v2 hierarchy is unified and different from v1's.
Cgroup driver: the detail that breaks silently
The subtlest point in the article, and probably the one that causes the most production incidents, is the alignment between the kubelet's cgroup driver and the container runtime's. The two need to match; if they don't, resource isolation behavior becomes inconsistent between what Kubernetes thinks it's applying and what the kernel actually enforces.
The recommendation from the post's own author is direct:
If you can pick either option, I recommend using the systemd driver.
Paco Xu, DaoCloud
The good news is that KEP-4033, which deals with automatic cgroup driver discovery via CRI, graduated to stable in version 1.34. With a runtime that implements the RuntimeConfig RPC (containerd 2.0+ or CRI-O 1.28+), the kubelet starts using the value reported by the runtime itself instead of relying on the manual cgroupDriver: systemd setting in the kubelet config file. Anyone still running an older runtime still needs to configure this by hand, and forgetting this step is the most common cause of a node that comes up but applies the wrong limit.
Memory, OOM, and PSI: what changes in observed behavior
The part that matters most to whoever is on SRE on-call is what changes in runtime behavior, not just configuration. Three points from the post stand out:
- Memory QoS (alpha, updated in 1.36): depends on the cgroup v2 memory controller.
memory.highdoes throttling, whilememory.minandmemory.lowprovide tiered protection whenmemoryReservationPolicy: TieredReservationis enabled. This simply doesn't exist in cgroup v1. The risk detail: on kernels older than 5.9,memory.highreclaim can get stuck in a known livelock, and 1.36 started logging a warning when it detects this combination. - OOM per container, not per process: on cgroup v2 nodes, the kubelet starts using
memory.oom.groupby default (viasingleProcessOOMKill: false), so an OOM event kills every process in that container together, instead of taking down a single process and leaving the container running half-broken. Anyone relying on the old behavior needs to explicitly setsingleProcessOOMKill: true. - PSI (Pressure Stall Information) in GA: the
KubeletPSIfeature is stable and enabled by default on clusters that meet the requirements (cgroup v2, kernel 4.20+,CONFIG_PSI=y, not booted withpsi=0). This is data exposed via the Summary API and/metrics/cadvisor, useful for telling real CPU/memory/IO contention apart from simply high usage.
None of these three features work on cgroup v1. That's the strongest argument for migrating beyond the imposed deadline: staying on v1 isn't just keeping the status quo, it's giving up visibility that's already available and that helps diagnose throttling and contention without extra instrumentation.
How to audit the cluster before upgrading
The post provides an end-to-end verification roadmap that's worth replicating in any pre-upgrade pipeline. At the OS level, stat -fc %T /sys/fs/cgroup/ returns cgroup2fs when the node is already on v2. To inspect the tree of units managed by systemd, systemctl list-units 'kube*' --type=slice lists the Kubernetes slices; tree -L 2 -d /sys/fs/cgroup/kubepods.slice shows the Pods' directory structure.
To confirm that a CPU or memory limit actually made it down to the cgroup, the documented path is to go layer by layer: compare spec.containers[*].resources with status.containerStatuses[*].resources via kubectl, locate the container with crictl inspect, extract the PID, and then read cpu.weight, cpu.max, and memory.max directly in the corresponding cgroup. It's manual work, but it's the only reliable way to confirm that an in-place resize (stable since 1.35, with Pod-level resource support in beta in 1.36) actually converged, since status can temporarily fall out of sync with spec during the adjustment.
What this costs, in practice
The migration itself has no licensing cost and requires no workload rewrite: applications don't need to know whether they're running on v1 or v2. The real cost is operational and concentrates on three fronts: upgrading the kernel to 5.8+ across fleets of nodes still running old distributions, validating that the entire containerd/CRI-O fleet is on minimum versions, and auditing any third-party tool (monitoring, policy engine, custom exporter) that reads the cgroup filesystem directly, because the hierarchy's structure has changed.
For those already on a recent kernel and updated runtime, the window until version 1.38 gives plenty of time to spare. For those still carrying legacy nodes in production, this is the deprecation worth putting on the radar for the next upgrade cycle, not leaving it for the day the kubelet simply refuses to start.
Translated from the Brazilian Portuguese original · Read the original
Pod Security Admission Is the Official Path to Retire PodSecurityPolicy in Kubernetes
PodSecurityPolicy has been deprecated by Kubernetes, but many clusters still run without any security admission control in its place. The official documentation details how Pod Security Admission solves this per namespace, without requiring an external webhook.