Karpenter Blueprint: Static NodePool¶
Purpose¶
The purpose of this blueprint is to demonstrate how to maintain a fixed number of nodes running in your cluster regardless of workload demand. Setting spec.replicas in a NodePool tells Karpenter to maintain a node count. This blueprint walks through setting up a static node pool for GPU instances alongside configuring network interfaces (ENA or EFA) and how to configure capacity reservations or placement groups.
You might consider this when: - You've purchased capacity via EC2 Capacity Blocks for ML or On-Demand Capacity Reservations and want a fixed node count - You need a static cluster of accelerated nodes with EFA networking for distributed training
Requirements¶
- A Kubernetes cluster with Karpenter installed. Karpenter v1.11+ is required for placement group and network interface configuration support. You can use the cluster we've used to test this pattern at the
clusterfolder in the root of this repository. - The
StaticCapacityfeature gate must be enabled. This feature is in Alpha (default:false) since Karpenter v1.8.x. Update your Karpenter deployment to include the feature gate:
helm registry logout public.ecr.aws
helm upgrade karpenter oci://public.ecr.aws/karpenter/karpenter \
--namespace karpenter \
--set "settings.featureGates.staticCapacity=true" \
--reuse-values
Alternatively, if you're using the Terraform template from this repository, you can add the feature gate to the Karpenter Helm values.
Verify the feature gate is enabled by checking the Karpenter deployment configuration:
kubectl -n karpenter get deployment karpenter -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="FEATURE_GATES")].value}'
You should see StaticCapacity=true in the output.
For accelerated instances, you will also need your cluster to have the required drivers and device plugins to advertise accelerators to Kubernetes. You can learn more about what is included as part of the EKS-optimized accelerated AMIs here.
Deploy¶
Before applying the manifests, set your cluster-specific variables. If you're using the Terraform template provided in this repo, run the following commands:
export CLUSTER_NAME=$(terraform -chdir="../../cluster/terraform" output -raw cluster_name)
export KARPENTER_NODE_IAM_ROLE_NAME=$(terraform -chdir="../../cluster/terraform" output -raw node_instance_role_name)
NOTE: If you're not using Terraform, you need to get those values manually.
CLUSTER_NAMEis the name of your EKS cluster (not the ARN). Karpenter auto-generates the instance profile in yourEC2NodeClassgiven the role that you specify in spec.role with the placeholderKARPENTER_NODE_IAM_ROLE_NAME, which is a way to pass a single IAM role to the EC2 instance launched by the KarpenterNodePool. Typically, the instance profile name is the same as the IAM role (not the ARN).
Create the EC2NodeClass¶
This defines the AWS-specific config for GPU nodes in the static pool:
cat << EOF > static-nodeclass.yaml
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
name: gpu-static
spec:
amiSelectorTerms:
- alias: bottlerocket@latest
role: "$KARPENTER_NODE_IAM_ROLE_NAME"
networkInterfaces:
- networkCardIndex: 0
deviceIndex: 0
interfaceType: "interface"
- networkCardIndex: 0
deviceIndex: 1
interfaceType: "efa-only"
securityGroupSelectorTerms:
- tags:
karpenter.sh/discovery: $CLUSTER_NAME
subnetSelectorTerms:
- tags:
karpenter.sh/discovery: $CLUSTER_NAME
instanceStorePolicy: RAID0
EOF
Network interface configuration (ENA / EFA)¶
You might decide to attach EFA devices for high-throughput inter-node communication (e.g. RDMA). To do this you can configure networkInterfaces as part of the EC2NodeClass. Two interface types are available:
interface— standard ENA providing IP connectivityefa-only— EFA device for RDMA, doesn't consume an IP address
The interface entry handles IP networking. The efa-only entry attaches an EFA device for RDMA traffic. Instances with multiple network cards (e.g., p6-b200.48xlarge) need additional entries.
The configuration in the EC2NodeClass above defines that instances launched by this EC2NodeClass primary network interface is ena and secondary as efa-only.
NOTE: Network interface configuration varies by instance type. Check the EFA documentation and EC2 network specifications for your specific instance.
Placement groups¶
To specify an EC2 placement group, add a placementGroupSelector to the EC2NodeClass.
placementGroupSelector:
name: my-gpu-cluster-pg
You might use placement groups to co-locate compute for performance, or spread for resilience. Karpenter supports three placement group strategies: - Cluster — single AZ, same network segment, best for EFA workloads - Partition — up to 7 isolated partitions per AZ for fault isolation - Spread — each instance on distinct hardware, max 7 per AZ per group
NOTE: The placement group must already exist before applying the EC2NodeClass — Karpenter does not create placement groups. Placement group support requires Karpenter v1.11.0+. See the Karpenter EC2NodeClass documentation for more information.
Placement group gotchas — apply the same way on both OSS Karpenter and EKS Auto Mode:
- Cluster PG pins the AZ on first launch. Once the first node lands, subsequent nodes in that PG must be in the same AZ. Pin
topology.kubernetes.io/zonein your NodePool requirements or expect racy scale-ups when a different AZ is chosen.- Spread PG caps at 7 instances per AZ per group. At the cap, Karpenter's drift replacement is blocked. Use
consolidationPolicy: WhenEmptyon the NodePool so an outgoing node is fully drained before a replacement is attempted.- Pods don't inherit PG membership. Consolidation can schedule pods onto nodes outside the PG. If the pods must live inside it, express that with a
nodeSelector(e.g. matching the NodePool label) or atopologySpreadConstraint.- Deleted PGs drift silently.
placementGroupSelectoris validated at admission, but the PG's continued existence is only checked at node-launch time. Delete a PG a NodeClass still points at and Karpenter will keep trying to use it, then fail at launch.
Capacity reservations¶
If using capacity via ODCRs or Capacity Blocks, add capacityReservationSelectorTerms to target it using id or tags:
capacityReservationSelectorTerms:
- tags:
karpenter.sh/discovery: ${CLUSTER_NAME}
- id: cr-123
Tagging your reservations at purchase time (e.g., karpenter.sh/discovery: ${CLUSTER_NAME}) makes them easier to manage as Karpenter can discover them automatically via tag selectors instead of requiring you to track individual reservation IDs. This is especially useful when you have multiple reservations.
To use reserved capacity in your NodePool, add reserved to the karpenter.sh/capacity-type requirement:
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["reserved"]
Create the NodePool¶
The NodePool maintains the static GPU node count of g6e.8xlarge.
cat << EOF > static-nodepool.yaml
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: gpu-static
spec:
replicas: 1
limits:
nodes: 2
template:
metadata:
labels:
capacity-type: gpu-static
nvidia.com/gpu.present: "true"
vpc.amazonaws.com/efa.present: "true"
spec:
nodeClassRef:
group: karpenter.k8s.aws
name: gpu-static
kind: EC2NodeClass
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
- key: node.kubernetes.io/instance-type
operator: In
values: ["g6e.8xlarge"]
taints:
- key: nvidia.com/gpu
effect: NoSchedule
EOF
In this example, the karpenter.sh/capacity-type was set to on-demand, to use reserved capacity (ODCRs or Capacity Blocks), add reserved to the karpenter.sh/capacity-type values and check the EC2NodeClass references the reserved capacity otherwise reservations will not be used. You need to set node.kubernetes.io/instance-type to the reserved instance type so it matches the capacity reservation.
NOTE: Adjust
replicas,limits.nodes, and instance type for your setup.limits.nodescaps the node count during scaling or drift replacement.
To specify an AZ for static capacity based on your reservation, you can add the topology.kubernetes.io/zone to your NodePool:
- key: topology.kubernetes.io/zone
operator: In
values: ["us-east-2a"] # Change to your target AZ
Apply the manifests¶
kubectl apply -f static-nodeclass.yaml
kubectl apply -f static-nodepool.yaml
Since replicas is set, Karpenter provisions the nodes without any pending pods.
Results¶
After a few minutes, Karpenter provisions a g6e.8xlarge matching the replicas: 1 spec.
Check nodes:
kubectl get nodes -l capacity-type=gpu-static -o custom-columns="NAME:.metadata.name,INSTANCE:.metadata.labels.node\.kubernetes\.io/instance-type,READY:.status.conditions[-1].status,AGE:.metadata.creationTimestamp"
Expected output:
NAME INSTANCE READY AGE
ip-10-0-x-x.region.compute.internal g6e.8xlarge True 2026-04-14T00:00:00Z
Check NodePool status:
kubectl get nodepool
Expected output:
NAME NODECLASS NODES READY AGE
gpu-static gpu-static 1 True 48s
...
EKS Auto Mode
**Prerequisite:** an EKS cluster with Auto Mode enabled, and an EKS Access Entry granting `AmazonEKSAutoNodePolicy` to the node IAM role used by Auto Mode. > If you're using the Terraform template under [`cluster/automode/`](https://github.com/aws-samples/karpenter-blueprints/tree/main/cluster/automode) in this repo, the cluster, node IAM role, and Access Entry are all created for you — you can skip the manual access entry steps below. EKS Auto Mode supports static capacity NodePools with the same `spec.replicas` field — see the [Static Capacity Node Pools in EKS Auto Mode](https://docs.aws.amazon.com/eks/latest/userguide/auto-static-capacity.html) documentation. **No `StaticCapacity` feature gate flip is needed** — the managed Karpenter in Auto Mode exposes this directly. As of the [July 2026 EFA and Placement Groups launch](https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-eks-efa-placement-groups/), the Auto Mode `NodeClass` also exposes EFA network interfaces and placement-group configuration as first-class fields. The shipped `static-nodeclass-automode.yaml` sets a primary `interface` ENI plus a secondary `efa-only` ENI under `spec.advancedNetworking.networkInterfaces`, and the paired `static-nodepool-automode.yaml` carries the `vpc.amazonaws.com/efa.present: "true"` label so pods that request EFA can schedule onto it. NVIDIA drivers and the device plugin are still bundled by Auto Mode, so no extra install is needed. To deploy on Auto Mode, set the variables from the Auto Mode Terraform template (not the OSS `cluster/terraform` one used earlier in this README):export CLUSTER_NAME=$(terraform -chdir="../../cluster/automode" output -raw cluster_name)
export KARPENTER_NODE_IAM_ROLE_NAME=$(terraform -chdir="../../cluster/automode" output -raw node_role_name)
sed -i \
-e "s/<<CLUSTER_NAME>>/$CLUSTER_NAME/g" \
-e "s/<<KARPENTER_NODE_IAM_ROLE_NAME>>/$KARPENTER_NODE_IAM_ROLE_NAME/g" \
static-nodeclass-automode.yaml
kubectl apply -f static-nodeclass-automode.yaml
kubectl apply -f static-nodepool-automode.yaml
kubectl get nodes -l capacity-type=gpu-static \
-o custom-columns="NODE:.metadata.name,EFA_LABEL:.metadata.labels.vpc\.amazonaws\.com/efa\.present,EFA_ALLOCATABLE:.status.allocatable.vpc\.amazonaws\.com/efa"
INSTANCE_ID=$(kubectl get nodes -l capacity-type=gpu-static -o jsonpath='{.items[0].metadata.name}')
aws ec2 describe-instances --instance-ids $INSTANCE_ID \
--query 'Reservations[0].Instances[0].{PlacementGroup:Placement.GroupName,ENIs:NetworkInterfaces[*].{DeviceIndex:Attachment.DeviceIndex,Type:InterfaceType,PrivateIp:PrivateIpAddress}}'
{
"PlacementGroup": "my-placement-group",
"ENIs": [
{ "DeviceIndex": 1, "Type": "efa-only", "PrivateIp": null },
{ "DeviceIndex": 0, "Type": "interface", "PrivateIp": "10.0.x.x" }
]
}
aws eks create-access-entry \
--cluster-name $CLUSTER_NAME \
--principal-arn <node-role-arn> \
--type EC2
aws eks associate-access-policy \
--cluster-name $CLUSTER_NAME \
--principal-arn <node-role-arn> \
--policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSAutoNodePolicy \
--access-scope type=cluster
Sample workloads¶
This blueprint provisions the node substrate — the actual workload is out of scope because it varies by model, dataset, and framework. Common fits for a static pool of accelerated nodes with EFA and a placement group are distributed training (PyTorch DDP or Megatron-style multi-node runs) and long-running model inference (vLLM, TGI) where cold-start churn on autoscaled pools would hurt latency.
Actually flexing the EFA fabric at runtime (rather than just declaring it in the NodeClass) is beyond what this blueprint covers — training runs are typically hours long and their setup lives with the ML sample, not with the compute plumbing. For a lightweight way to verify the fabric on its own, the NVIDIA NCCL tests repository is the standard benchmark; the AWS samples on GitHub have end-to-end distributed training examples that layer on top of a pool like the one this blueprint sets up.
Clean-up¶
To clean-up execute the following commands:
kubectl delete -f static-nodepool.yaml
kubectl delete -f static-nodeclass.yaml