3.7 NodeSelectors, Node Affinity & Anti-Affinity
Key Takeaways
nodeSelector provides simple key-value matching to bind Pods to labeled nodes, whereas Node Affinity provides expressive Boolean expressions and weighted preferences.
Node Affinity rules are bifurcated into hard constraints (requiredDuringSchedulingIgnoredDuringExecution) and soft preferences (preferredDuringSchedulingIgnoredDuringExecution).
Supported node affinity operators include In, NotIn, Exists, DoesNotExist, Gt, and Lt, evaluated exclusively at scheduling time.
Inter-Pod Affinity co-locates related services in the same topology domain, while Pod Anti-Affinity distributes replicas across distinct failure domains (nodes, zones, regions).
topologyKey defines the failure domain boundary (e.g., kubernetes.io/hostname, topology.kubernetes.io/zone) across which pod affinity and anti-affinity rules are evaluated.
Node Selectors, Node Affinity & Anti-Affinity
Default Kubernetes scheduling distributes Pods across nodes based on available capacity. However, real-world enterprise architectures require deterministic workload placement: co-locating web services with in-memory caches, binding GPU-intensive workloads to specialized hardware nodes, or spreading database replicas across distinct cloud availability zones to withstand data center failures.
1. Node Placement Mechanisms Hierarchy
Kubernetes provides multiple levels of node targeting and workload distribution:
+-----------------------------------------------------------------------------------------+
| NODE SELECTION COMPARISON |
| |
| 1. spec.nodeName |
| - Hardcoded node name (e.g., nodeName: "worker-01"). |
| - Completely BYPASSES kube-scheduler (Pod placed directly on node). |
| |
| 2. spec.nodeSelector |
| - Simple key-value equality matching against Node labels. |
| - Evaluated by kube-scheduler (e.g., disktype: "ssd"). |
| |
| 3. spec.affinity.nodeAffinity |
| - Expressive Boolean logic (In, NotIn, Exists, DoesNotExist, Gt, Lt). |
| - Supports Hard rules (required) and Soft weighted rules (preferred). |
| |
| 4. spec.affinity.podAffinity / podAntiAffinity |
| - Places Pods based on labels of OTHER RUNNING PODS in a topology domain. |
+-----------------------------------------------------------------------------------------+
# Label a node for nodeSelector matching
kubectl label nodes worker-node-02 disktype=ssd hardware=gpu
# Using simple nodeSelector
spec:
nodeSelector:
disktype: ssd
Caution
If no worker nodes match ALL labels in nodeSelector, the Pod remains stuck in Pending state with 0/N nodes available: node(s) didn't match Pod's node selector.
2. Node Affinity Deep Dive (spec.affinity.nodeAffinity)
Node Affinity extends nodeSelector with expressive query syntax and soft preference weighting.
Types of Node Affinity
requiredDuringSchedulingIgnoredDuringExecution(Hard Rule):- The scheduler must place the Pod on a node that satisfies the rules. If no matching node is found, the Pod remains in
Pendingstate.
- The scheduler must place the Pod on a node that satisfies the rules. If no matching node is found, the Pod remains in
preferredDuringSchedulingIgnoredDuringExecution(Soft Rule):- The scheduler attempts to place the Pod on matching nodes. If no matching nodes exist, the Pod is scheduled onto any available node.
- Each preference carries a
weightfrom1to100. The scheduler calculates a cumulative score for each candidate node by summing the weights of all satisfied preferences.
Note
IgnoredDuringExecution means that if a node's labels change after a Pod is scheduled (such that the affinity rule is no longer met), the Pod continues to run without being evicted.
apiVersion: apps/v1
kind: Deployment
metadata:
name: analytics-worker
spec:
replicas: 3
selector:
matchLabels:
app: analytics
template:
metadata:
labels:
app: analytics
spec:
affinity:
nodeAffinity:
# 1. Hard Rule: Must be in us-east-1a or us-east-1b AND have SSD
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values:
- us-east-1a
- us-east-1b
- key: disktype
operator: In
values:
- ssd
# 2. Soft Rule: Prefer GPU nodes if available
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 80
preference:
matchExpressions:
- key: accelerator
operator: In
values:
- nvidia-h100
containers:
- name: worker
image: python:3.11
Node Affinity Match Operators
In: Label value must match one of the specified strings invalues.NotIn: Label value must not match any of the specified strings (used for anti-affinity).Exists: Node must possess the label key (novaluesarray required).DoesNotExist: Node must not possess the label key.Gt/Lt: Integer comparison (label value parsed as integer).
3. Inter-Pod Affinity and Anti-Affinity
While Node Affinity targets node labels, Pod Affinity and Anti-Affinity evaluate the labels of other Pods already running on candidate nodes within a specific failure domain (topologyKey).
+-----------------------------------------------------------------------------------------+
| POD AFFINITY VS POD ANTI-AFFINITY |
| |
| 1. POD AFFINITY (Co-location) |
| "Schedule Web-App Pods in the same zone where Redis-Cache Pods reside." |
| |
| 2. POD ANTI-AFFINITY (Distribution / High Availability) |
| "DO NOT schedule two Web-App Pods onto the same physical node (topologyKey: host).|
| "DO NOT schedule two Database Pods in the same Cloud Zone (topologyKey: zone)." |
+-----------------------------------------------------------------------------------------+
apiVersion: apps/v1
kind: Deployment
metadata:
name: ha-web-service
spec:
replicas: 3
selector:
matchLabels:
app: web
template:
metadata:
labels:
app: web
spec:
affinity:
# Anti-Affinity: Spread across physical nodes
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app
operator: In
values:
- web
topologyKey: "kubernetes.io/hostname"
# Affinity: Co-locate with in-memory redis cache in same zone
podAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app
operator: In
values:
- redis
topologyKey: "topology.kubernetes.io/zone"
containers:
- name: nginx
image: nginx:alpine
Understanding topologyKey
The topologyKey defines the geographic or physical boundary of the domain:
kubernetes.io/hostname: Per-node domain (ensures Pods are on different physical servers).topology.kubernetes.io/zone: Availability zone domain (ensures Pods are in different data centers).topology.kubernetes.io/region: Cloud region domain.
4. Topology Spread Constraints (spec.topologySpreadConstraints)
In addition to binary affinity rules, Kubernetes supports Topology Spread Constraints to achieve an even distribution of Pods across topology domains:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: web
maxSkew: Maximum allowable difference in Pod count between any two topology domains.whenUnsatisfiable: DoNotSchedule(Hard) orScheduleAnyway(Soft).
5. Diagnostic Commands and Troubleshooting Placement Failures
# Identify why a Pod is stuck in Pending due to Affinity
kubectl describe pod <pod-name>
# Events: 0/3 nodes are available: 3 node(s) didn't match Pod topology spread constraints / pod anti-affinity.
# Check node labels across the cluster
kubectl get nodes --show-labels
# Filter nodes matching specific label selector
kubectl get nodes -l topology.kubernetes.io/zone=us-east-1a
# Imperatively add or remove node labels
kubectl label nodes worker01 hardware=gpu --overwrite
kubectl label nodes worker01 hardware- # Remove label
An administrator wishes to ensure that Pods for an analytics application only run on nodes in availability zones 'zone-1' or 'zone-2', and that the Pods MUST NOT be scheduled anywhere else. Which configuration is required?
spec.nodeName: [zone-1, zone-2]
spec.affinity.nodeAffinity.preferredDuringSchedulingIgnoredDuringExecution with weight: 100
spec.affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution with matchExpressions using operator: In and values: [zone-1, zone-2]
spec.tolerations with effect: NoSchedule and key: zone
What is the primary function of the 'topologyKey' field in Pod Anti-Affinity rules?
It defines the encryption cipher key used for node-to-node Pod networking.
It specifies the maximum number of network hops between communicating microservices.
It sets the storage class topology identifier for dynamic volume provisioning.
It defines the node label key that demarcates the domain (e.g. node, zone, rack) across which Pod co-location or separation is enforced.
A cluster has 3 worker nodes. An administrator deploys a Deployment with replicas: 5 and configures a hard podAntiAffinity rule with topologyKey: 'kubernetes.io/hostname' matching the Deployment's own label selector. What will be the state of the 5 Pods?
3 Pods will be Running (one on each node) and 2 Pods will remain in Pending state because no additional distinct nodes are available to satisfy the hard anti-affinity rule.
All 5 Pods will run successfully because Kubernetes automatically overrides anti-affinity when replicas exceed node count.
All 5 Pods will fail with CrashLoopBackOff due to a port conflict on the host network.
2 Pods will be scheduled on the control plane node without tolerations.
Sections you finish are checked off in the contents.