Open Issues Need Help
View All on GitHub Notes: Handling LLM / Agent Access to OpenShift Clusters about 1 month ago
documentation enhancement help wanted research openshift rhoai observability coldfront NAIRR MOC 2.0
Explore & Optimize: Ensuring Effective Log Storage Management for the Observability Cluster about 1 month ago
documentation enhancement help wanted research openshift observability needs_clarification ope MOC 1.0
Explore & Connect: Making long term logs (archive, S3) viewable in OBS logging console about 1 month ago
enhancement help wanted openshift observability MOC 1.0
Discussion: GPU Node Toleration Policy - Should pods without GPU requests be allowed on GPU nodes? about 1 month ago
bug documentation enhancement help wanted openshift needs_clarification gpu H100 Optimization NAIRR MOC 1.0
Cleanup RHOAI Jupyter Notebook Image list about 2 months ago
enhancement help wanted rhoai MOC 1.0
Lag and timeout issue 4 months ago
help wanted ope
bug help wanted openshift ope
AI Summary: The `wrk-2` node in the OBS cluster has been unreachable since January 7, 2026, causing the AI Telemetry service to go down as its Kafka and Zookeeper pods were scheduled on it. This issue has also degraded the `machine-config` ClusterOperator and affected 104 other pods. A hardware check is recommended to resolve the complete node outage.
Complexity:
4/5
bug help wanted openshift observability ai-telemetry