Resource state, quota accounting and measured utilization are different signals. Kueue reports request-based quota accounting; JobSet member views can show current CPU/memory observations. Neither establishes physical GPU capacity or measured GPU utilization, and this ecosystem coverage is not a complete scheduling simulation.
Queueing & scheduling
Workload status follows the Kueue admission lifecycle - Pending, QuotaReserved, Admitted, Evicted/Preempted, Finished. Queue badges reflect
Active (Kueue) or Open/Closed (Volcano) state.
Kueue admission and queue detail
Start with a pending Workload to inspect its PodSets, controller-evaluated resource requests, quota reservation, assigned flavors, admission-check states and messages, and native conditions. Follow its submission LocalQueue to the backing ClusterQueue even before a reservation exists. ResourceFlavor and AdmissionCheck references also open their resources. You can also start from a batch Job or JobSet. Drawer and fullscreen details show exact controller-owned Kueue Workloads, admission decisions, native reasons, conditions and queue links separately from execution status. Ordinary Jobs without Kueue hints or observed admission evidence stay quiet. Suspension alone does not establish an admission blocker, and a parent-owned Workload is not attributed to a child Job. Unavailable observations and list limits are explicit. LocalQueue and ClusterQueue detail supportkueue.x-k8s.io/v1beta1 and v1beta2:
- State: reported Active condition and message, configured stop policy, and pending/reserving/admitted counts. Intentional stops remain distinct from failures; stale conditions and missing values are explicit.
- Quota: reservations and admitted usage by flavor and resource. ClusterQueues additionally show nominal quota and applicable borrowing/lending limits, preserving explicit zero limits.
- Policy: namespace eligibility, cohort, queueing/preemption settings and required admission checks, including flavor applicability.
kueue.x-k8s.io/v1beta2 Workloads, REST AI detail and MCP get_resource also include structured scheduling context with controller evidence and queue roles. Queue position, start-time predictions, queue-wide Workload lists and historical utilization are not provided by these detail views.
Volcano’s Job shares its kind name with the built-in batch/v1 Job, and Volcano and KAI both ship Queue and PodGroup kinds. Radar disambiguates by API group everywhere - tables, filters, and status badges always route to the right tool.
Distributed training & batch
Training jobs show per-replica-type readiness (Master 1/1, Worker 3/4) and elapsed time; Ray kinds surface job status, deployment status, and worker counts.
JobSet investigation
Forjobset.x-k8s.io/v1alpha2, inspect root lifecycle, per-role observations, dependencies, completion/restart policies and controller conditions. Fullscreen detail connects the JobSet to its controller-owned member Jobs and associated Kueue Workloads, keeping admission evidence separate from execution state.
Filter retained member Jobs by role, name or state, then select a Job to inspect its Pods and live logs. Role-wide and whole-JobSet logs are bounded snapshots with source attribution and read failures. Current CPU/memory comparisons and totals report their observation coverage; missing or stale samples are not zeros. Extended-resource requests describe declared demand, not GPU/TPU utilization.
REST AI detail and MCP get_resource include structured execution context. A failed child does not by itself establish a terminal JobSet failure, and missing children do not imply success. Unavailable observations and list limits are explicit. Historical goodput and archived logs are not collected.
RayService context for AI clients
Forray.io/v1 RayServices, REST AI detail and MCP get_resource distinguish active and pending revisions, application states, upgrade/rollback evidence and suspension. A healthy active revision can coexist with a failing pending revision. Reported traffic percentages are configured route weights, not measured requests. This structured context does not imply a dedicated RayService dashboard or live traffic measurement.
Inference serving
InferenceService badges distinguish model-load failures (shown as Load failed, from the
BlockedByFailedLoad transition status) from plain not-ready; InferencePool acceptance reflects per-gateway Accepted + ResolvedRefs conditions across both API groups.
GPU operators
The NVIDIA GPU Operator and DRA get full typed detail views - see their dedicated pages. Classic extended-resource GPU visibility (node capacity, pod requests, GPU table columns) works on any cluster with no operator at all.