Kubernetes v1.38.0 is scheduled for 16 December 2026, and the kubelet in that release stops falling back to the cgroupDriver field in KubeletConfiguration when the container runtime cannot answer the CRI RuntimeConfig RPC. containerd 1.x never implemented that RPC, the backport was closed without merging, and containerd 1.7 reached end of life on 30 September 2026. A node in that state does not degrade gracefully: the kubelet refuses to start and the node never registers.

The exposure is invisible from a dashboard. Every affected node is Ready right now, serving production, reporting healthy, and nothing in the control plane will hint at it until someone rolls a 1.38 kubelet. The behaviour change lives in kubernetes/kubernetes#136731, still open as of 2 October 2026, so most current documentation describes the v1.34 GA behaviour where the fallback still exists. These checks are for whoever owns the node estate and has to answer "are we exposed" before the December window opens. They sit alongside two other reasons a node stops joining in this release train: the 18 kubelet flags that became fatal in 1.37 and cgroup v1 on a v2 host, both of which fail at a different point in startup.

kubelet startsCalls CRI RuntimeConfig RPCRuntime returnsa cgroup driver?Use the driverthe runtime reportsDisableCgroupDriverFallbackon? (default in 1.38)kubelet refuses to start,node never joinsFall back to cgroupDriverin KubeletConfiguration"yes""no""yes""no"

The checks

  1. Inventory from the field the kubelet publishes, not from your configuration management. The authoritative value is status.nodeInfo.containerRuntimeVersion on every Node object, and it reports what is actually running, including nodes built from an AMI nobody can account for:
   kubectl get nodes -o json \
     | jq -r '.items[] | [.metadata.name, .status.nodeInfo.containerRuntimeVersion] | @tsv' \
     | awk '{print $2}' | sort | uniq -c | sort -rn

Anything printing containerd://1. is a node that will not come back from a 1.38 kubelet.

  1. Treat the version string as evidence, then ask the runtime the question the kubelet asks. crictl ships a runtime-config subcommand that retrieves the runtime configuration over the same RPC, so it answers directly instead of by inference:
   crictl -r unix:///run/containerd/containerd.sock runtime-config

A runtime that implements the RPC prints a cgroup driver. A runtime that does not returns a gRPC error for an unimplemented method. Record the output per node pool rather than per cluster, and count anything that is not a printed driver as a failing node.

  1. Canary the removal on a cordoned node this week, before any 1.38 binary exists in your estate. Cgroup driver autoconfiguration went GA in v1.34, so on a supported runtime the CRI answer already wins today. Cordon one node per pool, remove the cgroupDriver line from /var/lib/kubelet/config.yaml, restart the kubelet, and read the shape of the pod cgroup paths: kubepods-burstable-pod<uid>.slice means systemd is in effect, a /sys/fs/cgroup/kubepods/burstable/ tree means cgroupfs. A node that keeps systemd with the field absent is not using the fallback at all and will survive December untouched. Keep it cordoned until you have read the result, because an unsupported runtime silently drops that node to cgroupfs.
  2. Stop planning around a containerd 1.7 patch. The backport that would have given 1.7 the RuntimeConfig RPC, containerd/containerd#11346, was closed without merging on 10 April 2025 after a community meeting. Its own description states the consequence: "When the feature is GA'd containerd v1.7 becomes incompatible with kubernetes. The kubelet refuses to start." containerd's RELEASES.md lists the 1.7 branch as end of life since 30 September 2026. The only path off this is a runtime upgrade on every affected node.
  3. Know the escape hatch, and write its expiry date into the same ticket. PR #136731 introduces a feature gate named DisableCgroupDriverFallback, enabled by default in 1.38 and deprecated on arrival. Setting it to false restores the old behaviour while you finish the runtime rollout:
   # /var/lib/kubelet/config.yaml
   apiVersion: kubelet.config.k8s.io/v1beta1
   kind: KubeletConfiguration
   featureGates:
     DisableCgroupDriverFallback: false

Two caveats worth five seconds each. The PR is still open, so confirm the gate name against the 1.38 changelog once code freeze passes on 16 November 2026 per the release-1.38 schedule. And budget for the gate disappearing in 1.39 or 1.40, which rules it out as a rollback plan.

  1. After the runtime upgrade, a different file decides your cgroup driver. Once containerd answers the RPC, /etc/containerd/config.toml is authoritative and the kubelet's own setting is ignored. containerd resolves the value from the first runtime handler that declares one, starting with the default handler, and auto-detects only when no handler says anything. A containerd config default run during the upgrade resets SystemdCgroup to false, and that single line moves every pod on the node to cgroupfs while the host runs systemd:
   containerd config dump | grep -i SystemdCgroup
  1. Read the effective config, never the file on disk. The key path moved between branches: containerd 1.x used version = 2 with plugins."io.containerd.grpc.v1.cri", containerd 2.x uses version = 3 with plugins.'io.containerd.cri.v1.runtime'. A hand-maintained config.toml that keeps the old path parses without complaint while your SystemdCgroup = true sits where nothing reads it. Run containerd config migrate as part of the upgrade, then confirm with config dump.
  2. Count the nodes that do not exist yet. kubectl get nodes cannot see a scaled-to-zero pool, a spot pool between interruptions, a Windows pool, or the AMI pinned in a launch template that will hydrate the next node. A 1.7 runtime hides longest in those places, because the image was baked once and the pipeline that built it has not run since. Diff the AMI or image ID in every launch template, Karpenter node class and Cluster API machine template against the runtime version you verified in check 2.
  3. A managed node image is not automatically a safe one. AWS announced the move of EKS-optimized AL2023 AMIs from containerd 1.7.X to 2.1.X in amazon-eks-ami#2470, completed per Kubernetes version from release v20250915 onward, with 1.29 and 1.28 staying on containerd 1.7 until their EOL. That clears the 1.38 gate and lands you on containerd 2.1, which containerd's own release table marks end of life as of 3 July 2026. If you are choosing a version by hand, 2.0 is LTS to March 2027 and 2.3 is LTS to April 2028, while 2.2's window closes on 6 November 2026, before your upgrade even starts.
  4. Make the exposure a standing alert, not a spreadsheet that goes stale in a week. kube-state-metrics exports the runtime as a label on kube_node_info, so the fleet query is one expression:
    count(kube_node_info{container_runtime_version=~"containerd://1\\..*"}) > 0

KEP-4033 also commits the kubelet to logging a deprecation message about upgrading to a CRI implementation that supports cgroup driver detection. Confirm what your specific build prints to the journal before alerting on the text, and keep the metric as the signal you trust.

  1. Do not let this merge into the other node failures in the same window. Three separate changes can stop a node from joining across 1.35 to 1.38: cgroup v1 removal, the 1.37 kubelet flag removals, and the RuntimeConfig requirement. Each fails at its own point in startup and each has its own separating test, so one combined "upgrade readiness" ticket produces a pass that proves nothing. Keep them as three checks with three owners, alongside the SELinuxMount default that stalls shared PVCs, which bites after the node is already up.

Wrap-up

The habit worth keeping is check 2 as a standing pre-flight: crictl runtime-config against one node of every pool, every cycle, recorded per pool. The runtime upgrade itself belongs in its own change window rather than as a line item inside the Kubernetes upgrade, and decoupling the two is what buys you 1.36 and 1.37 to do this at a safe pace instead of during a December freeze.

The honest cost of the next step is a drain and reboot on every node, multiplied by the pools that cannot drain cleanly: StatefulSets with strict PodDisruptionBudgets, single-replica workloads nobody has re-platformed, and golden images whose build pipeline has not run in a year. Those are the hours, not the package install. The same drain budget is already claimed by kube-proxy moving from IPVS to nftables in 1.40, so scope this per pool now, in the right order, with a rollback that does not depend on a feature gate due to be deleted.