Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 25 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -317,6 +317,30 @@
<code>deserialize()</code> rather than reimplementing it.
</td>
</tr>
<tr>
<td><b>[Kubernetes] Add Intel GPU discovery and resource selection</b></td>
<td><code>5420c46</code></td>
<td>
<code>sky/provision/kubernetes/utils.py</code><br>
<code>sky/clouds/kubernetes.py</code><br>
<code>docs/source/reference/kubernetes/intel-gpu.rst</code><br>
<code>docs/source/compute/gpus.rst</code>
</td>
<td>
Adds Intel GPU discovery and counting through
<code>gpu.intel.com/xe</code> (<code>xe</code> driver only;
<code>i915</code> GPUs are not supported). The new
<code>IntelGPULabelFormatter</code> reads NFD
<code>gpu.intel.com/product</code> labels and normalizes model names
(e.g. <code>Arc_Pro_B60</code> → <code>Intel-Arc-Pro-B60</code>);
B-series product labels need Intel device plugin NFD rules v0.37.0+.<br>
Pods for Intel product labels request <code>gpu.intel.com/xe</code>,
alongside the existing AMD/NVIDIA label-key selection, so mixed
clusters and autoscaling without an existing GPU node work.<br>
Adds an Intel GPU setup and manual verification guide and links it
from the GPU documentation.
</td>
</tr>
</tbody>
</table>

Expand Down Expand Up @@ -396,6 +420,7 @@ git cherry-pick c6f8f23 # Raise per-controller service capacity for k8s
git cherry-pick 78fe751 # Pin uv pip to runtime venv via --python
git cherry-pick 69b0a69 # Exclude kubernetes==36.0.0 (in-cluster auth regression)
git cherry-pick 827ff42 # Accept PEP 585 dict[K,V] type strings in pod_config validator
git cherry-pick 5420c46 # Intel xe GPU discovery and resource selection
# Resolve any conflicts if upstream changed the same files

# 4. Push new branch
Expand Down
6 changes: 6 additions & 0 deletions docs/source/compute/gpus.rst
Original file line number Diff line number Diff line change
Expand Up @@ -95,9 +95,15 @@ AMD GPUs

See :ref:`kubernetes-amd-gpu`.

Intel GPUs
----------

See :ref:`kubernetes-intel-gpu`.

.. toctree::
:maxdepth: 1
:hidden:

Using Google TPUs <../../reference/tpu>
Using AMD GPUs <../../reference/kubernetes/amd-gpu>
Using Intel GPUs <../../reference/kubernetes/intel-gpu>
65 changes: 65 additions & 0 deletions docs/source/reference/kubernetes/intel-gpu.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
.. _kubernetes-intel-gpu:

Using Intel GPUs on Kubernetes
==============================

SkyPilot supports Intel GPUs that use the ``xe`` kernel driver and are exposed
through the ``gpu.intel.com/xe`` resource, such as Arc Pro B-series cards.
GPUs using the ``i915`` driver (``gpu.intel.com/i915``), including Arc
A-series, Data Center GPU Flex and Max, are not supported. Monitoring
resources are not counted as GPUs.

Cluster setup
-------------

Install the host ``xe`` GPU driver, the `Intel GPU device plugin
<https://intel.github.io/intel-device-plugins-for-kubernetes/cmd/gpu_plugin/README.html>`_,
and its Node Feature Discovery (NFD) rules, v0.37.0 or later. Use
``shared-dev-num=1`` if you want reported counts to correspond to unshared GPU
devices. With sharing enabled, Kubernetes advertises allocation slots rather
than physical GPU counts.

Each node must expose a single GPU model and a single GPU resource type.

GPU labels
----------

SkyPilot identifies Intel GPUs from the NFD ``gpu.intel.com/product`` label,
which the NFD rules set automatically:

.. code-block:: text

gpu.intel.com/product=Arc_Pro_B60 -> Intel-Arc-Pro-B60
gpu.intel.com/product=Arc_B580 -> Intel-Arc-B580

Request the GPU by its SkyPilot name, e.g. ``--gpus Intel-Arc-Pro-B60:1``.
Pods requesting an Intel GPU use the ``gpu.intel.com/xe`` resource. Because the
resource is derived from the label, this also works with autoscalers that
create Intel nodes on demand.

Workloads need a container image with the Intel userspace drivers and compute
runtime.

Manual verification
-------------------

After installing this version, restart the API server to pick up the changes:

.. code-block:: bash

sky api stop
sky api start
kubectl get nodes -o json
sky check kubernetes
sky show-gpus --infra kubernetes

Verify that each Intel node advertises ``gpu.intel.com/xe`` and has a
``gpu.intel.com/product`` label. The GPU listing should show the corresponding
model and allocatable count. Repeat with NVIDIA and AMD nodes in the same
cluster and confirm their counts are unchanged.

For a launch using an Intel-compatible image and the discovered accelerator,
inspect the resulting pod with ``kubectl get pod <pod-name> -o yaml``. Its GPU
request and limit should use ``gpu.intel.com/xe``. Check dashboard free counts
before and during the workload: one allocated GPU should reduce availability
by one.
7 changes: 6 additions & 1 deletion sky/clouds/kubernetes.py
Original file line number Diff line number Diff line change
Expand Up @@ -588,7 +588,8 @@ def _get_image_id(resources: 'resources_lib.Resources') -> str:
k8s_resource_key = kubernetes_utils.TPU_RESOURCE_KEY
else:
# Derive resource key from the matched label key.
# AMD device plugin labels start with 'amd.com/'; all other
# AMD device plugin labels start with 'amd.com/'; Intel NFD
# product labels map to the Xe resource; all other
# recognized GPU label formatters (GFD, SkyPilot, GKE,
# Karpenter, CoreWeave, Nebius) are for NVIDIA GPUs.
# We must NOT fall back to get_gpu_resource_key(context) here:
Expand All @@ -599,6 +600,10 @@ def _get_image_id(resources: 'resources_lib.Resources') -> str:
k8s_acc_label_key.startswith('amd.com/')):
k8s_resource_key = (
kubernetes_utils.SUPPORTED_GPU_RESOURCE_KEYS['amd'])
elif (k8s_acc_label_key
== kubernetes_utils.IntelGPULabelFormatter.LABEL_KEY):
k8s_resource_key = (kubernetes_utils.
SUPPORTED_GPU_RESOURCE_KEYS['intel_xe'])
elif k8s_acc_label_key is not None:
k8s_resource_key = (
kubernetes_utils.SUPPORTED_GPU_RESOURCE_KEYS['nvidia'])
Expand Down
67 changes: 52 additions & 15 deletions sky/provision/kubernetes/utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -172,16 +172,19 @@ def requires_tcpxo_daemon(self) -> bool:
'P': 2**50,
}

# The resource keys used by Kubernetes to track NVIDIA GPUs and Google TPUs on
# The resource keys used by Kubernetes to track GPUs and Google TPUs on
# nodes. These keys are typically used in the node's status.allocatable
# or status.capacity fields to indicate the available resources on the node.
SUPPORTED_GPU_RESOURCE_KEYS = {'amd': 'amd.com/gpu', 'nvidia': 'nvidia.com/gpu'}
SUPPORTED_GPU_RESOURCE_KEYS = {
'amd': 'amd.com/gpu',
'nvidia': 'nvidia.com/gpu',
'intel_xe': 'gpu.intel.com/xe',
}
TPU_RESOURCE_KEY = 'google.com/tpu'

NO_ACCELERATOR_HELP_MESSAGE = (
'If your cluster contains GPUs or TPUs, make sure '
f'one of {SUPPORTED_GPU_RESOURCE_KEYS["amd"]}, '
f'{SUPPORTED_GPU_RESOURCE_KEYS["nvidia"]} or '
f'one of {", ".join(SUPPORTED_GPU_RESOURCE_KEYS.values())} or '
f'{TPU_RESOURCE_KEY} resource is available '
'on the nodes and the node labels for identifying GPUs/TPUs '
'(e.g., skypilot.co/accelerator) are setup correctly. ')
Expand Down Expand Up @@ -791,6 +794,43 @@ def _normalize(cls, raw: str) -> str:
return name.replace(' ', '')


class IntelGPULabelFormatter(GPULabelFormatter):
"""Intel NFD product labels on nodes exposing a single GPU model.

e.g. gpu.intel.com/product=Arc_Pro_B60 -> Intel-Arc-Pro-B60. Only nodes
exposing the gpu.intel.com/xe resource are supported.
"""

LABEL_KEY = 'gpu.intel.com/product'

@classmethod
def match_label_key(cls, label_key: str) -> bool:
return label_key == cls.LABEL_KEY

@classmethod
def get_label_key(cls, accelerator: Optional[str] = None) -> str:
return cls.LABEL_KEY

@classmethod
def get_label_keys(cls) -> List[str]:
return [cls.LABEL_KEY]

@classmethod
def get_label_values(cls, accelerator: str) -> List[str]:
raise NotImplementedError

@classmethod
def validate_label_value(cls, value: str) -> Tuple[bool, str]:
valid = bool(value and value.strip())
return valid, (f'Intel GPU label value {value!r} must be non-empty.'
if not valid else '')

@classmethod
def get_accelerator_from_label_value(cls, value: str) -> str:
name = value.strip().replace('_', '-').replace(' ', '-')
return f'Intel-{name}'


def _accelerator_name_matches(requested_acc: str,
viable_names: List[str]) -> bool:
"""Check if requested accelerator matches any viable name.
Expand Down Expand Up @@ -882,8 +922,8 @@ def validate_label_value(cls, value: str) -> Tuple[bool, str]:
# auto-detecting the GPU label type.
LABEL_FORMATTER_REGISTRY = [
SkyPilotLabelFormatter, GKELabelFormatter, KarpenterLabelFormatter,
GFDLabelFormatter, AMDGPULabelFormatter, CoreWeaveLabelFormatter,
NebiusLabelFormatter
GFDLabelFormatter, AMDGPULabelFormatter, IntelGPULabelFormatter,
CoreWeaveLabelFormatter, NebiusLabelFormatter
]


Expand Down Expand Up @@ -1360,10 +1400,8 @@ def detect_accelerator_resource(
context: Optional[str]) -> Tuple[bool, Set[str]]:
"""Checks if the Kubernetes cluster has GPU/TPU resource.

Three types of accelerator resources are available which are each checked
with amd.com/gpu, nvidia.com/gpu and google.com/tpu. If amd.com/gpu or nvidia.com/gpu resource is
missing, that typically means that the Kubernetes cluster does not have
GPUs or the amd/nvidia GPU operator and/or device drivers are not installed.
Checks all supported GPU resources and Google TPUs. Missing resources
typically mean the device plugin or drivers are not installed.

Returns:
bool: True if the cluster has GPU_RESOURCE_KEY or TPU_RESOURCE_KEY
Expand Down Expand Up @@ -2058,15 +2096,14 @@ def get_accelerator_label_key_values(
'and re-run '
f'`sky ssh up --infra {context_display_name}`. {suffix}')
else:
gpu_resources = ', '.join(SUPPORTED_GPU_RESOURCE_KEYS.values())
msg = (
f'Could not detect GPU/TPU resources ({SUPPORTED_GPU_RESOURCE_KEYS["amd"]!r}, '
f'{SUPPORTED_GPU_RESOURCE_KEYS["nvidia"]!r} or '
f'Could not detect GPU/TPU resources ({gpu_resources} or '
f'{TPU_RESOURCE_KEY!r}) in Kubernetes cluster. If this cluster'
' contains GPUs, please ensure GPU drivers are installed on '
'the node. Check if the GPUs are setup correctly by running '
'`kubectl describe nodes` and looking for the '
f'{SUPPORTED_GPU_RESOURCE_KEYS["amd"]!r}, '
f'{SUPPORTED_GPU_RESOURCE_KEYS["nvidia"]!r} or '
f'{gpu_resources} or '
f'{TPU_RESOURCE_KEY!r} resource. '
'Please refer to the documentation on how to set up GPUs.'
f'{suffix}')
Expand Down Expand Up @@ -3812,7 +3849,7 @@ def get_node_accelerator_count(context: Optional[str],
gpu_resource_name = get_gpu_resource_key(context)
assert not (gpu_resource_name in attribute_dict and
TPU_RESOURCE_KEY in attribute_dict)
for gpu_resource in ["nvidia.com/gpu", "amd.com/gpu"]:
for gpu_resource in SUPPORTED_GPU_RESOURCE_KEYS.values():
if gpu_resource in attribute_dict:
return int(attribute_dict[gpu_resource])
if TPU_RESOURCE_KEY in attribute_dict:
Expand Down