Jobs · jobstreet_sg
AI Infrastructure Engineer
The Supreme HR Advisory Pte Ltd · Kaki Bukit, East Region · Posted 10d ago
About the role
AI Infrastructure Engineer 5 days, Mon - Fri 8.30am to 5.30pm Salary: $5,000 to $7,000 Location: 2 Kaki Bukit Ave 1, Singapore 417938 Job scopes: Compute & Cluster Management Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures). Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads. Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks. High-Performance Networking & Storage Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies). Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines. Automation & Infrastructure as Code (IaC) Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi. Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates. Operations, Observability & Performance Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface). Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs. Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments. Requirements: Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python). Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies. Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray). High-Speed Networking: Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations. Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints. Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible). Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience. 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure. Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect). If you are keen to apply, please send me your resume and job applied for at 9789 3505 or email me at [email protected] (˶ᵔ ᵕ ᵔ˶) ❄️Anabel Boon Xue Qi | Recruitment Consultant (R25159272) |📍The Supreme HR Advisory EA No: 14C7279
Read the full posting on jobstreet_sg →
FAQ
Is the AI Infrastructure Engineer role at The Supreme HR Advisory Pte Ltd remote?+
This AI Infrastructure Engineer position is listed as unknown (Kaki Bukit, East Region).
What is the salary for the AI Infrastructure Engineer role at The Supreme HR Advisory Pte Ltd?+
The listing states $5,000 – $7,000 per month.
What seniority level is this AI Infrastructure Engineer role?+
This is a unknown level position.
How do I apply for the AI Infrastructure Engineer role at The Supreme HR Advisory Pte Ltd?+
Use the "Apply on jobstreet_sg" button to open the original posting on jobstreet_sg, where you can submit your application directly to The Supreme HR Advisory Pte Ltd.