Platform Support Architect
DataDirect Networks
As an AI platform specialist, you will enable a cohesive HyperPOD AI data solution by aligning storage, AI software, and networking. You act as the primary NVIDIA AI Enterprise and vector DB expert, guiding design, diagnostics, and performance optimization for RAG workflows. You’ll develop repeatable runbooks, PoCs, and assets, while collaborating with NVIDIA, OEMs, and PS teams to translate field experiences into robust platform patterns. This role combines solutions architecture with engineering discipline to scale enterprise AI deployments.
Responsibilities- Serve as the primary NVIDIA AI Enterprise and Milvus vector DB expert for HyperPOD environments to guide diagnosis and solution design
- Own end-to-end triage across GPU, NVAIE services, vector DB, Kubernetes, Docker, high-speed networking, and Infinia storage to distinguish defects from environment issues
- Diagnose and resolve performance bottlenecks in RAG and agentic AI workflows, including model selection, prompts, and vector search
- Collect logs and telemetry across Linux, containers, Kubernetes, GPU stack, vector DB, and storage/networking; produce repros and defect reports for escalation
- Create and maintain Runbooks and triage checklists for HyperPOD covering NVAIE, Milvus, GPU stack, Kubernetes, and networking interactions
- Define unified diagnostics bundles across layers (Infinia, GPUs, NVAIE, Milvus, Kubernetes, network) to enable fast issue isolation
- Collaborate on observability tools to shape dashboards (Prometheus/Grafana/ELK/NetQ) that surface platform health and RAG SLIs
- Build hands-on labs/PoCs reflecting customer RAG/agentic AI use cases on HyperPOD to validate supportability and learnable configurations
- Develop reusable assets (guides, best-practice playbooks, tuning checklists, reference architectures) to accelerate customer value and support readiness
- Provide feedback to PM/Engineering on stack compatibility, upgrade/rollback constraints, and observability needs for NVAIE, Milvus/cuVS, Infinia, and networking
- Coordinate with NVIDIA and OEM architects, PS, and Support Innovation to align reference architectures with real-world support experience
- 5+ years in Linux-based infrastructure roles (SRE, MLOps, platform engineering, or L2/L3 support)
- 8+ years total technical experience preferred
- Strong hands-on experience with containers and Kubernetes (Docker/containerd, Helm, Operators)
- Proven track record operating GPU-accelerated workloads in production (NVIDIA GPUs, CUDA, GPU utilization triage)
- Familiarity with DGX/HGX or similar GPU clusters
- Experience with high-performance storage and HPC/AI networking (EXAScaler/Lustre, GPFS, Ceph; RDMA/InfiniBand)
- Hybrid cloud or cloud-adjacent patterns (Kubernetes CSI, cloud-native fabrics)
- Experience with vector databases (Milvus, Qdrant, Pinecone, pgVector, OpenSearch/Elasticsearch vectors)
- Solid understanding of RAG and Generative AI workflows and how they relate to vector search and GPU inference
- Familiarity with NVIDIA AI Enterprise components (NIM, NeMo, NeMo Retriever/Curator, Triton, TensorRT, CUDA libraries)
- Experience with ML/Ops pipelines: CI/CD for models, deployment strategies, canary/rollback, monitoring/alerting for AI services
- Strong diagnostic skills across Linux, containers, Kubernetes, GPUs, storage, and networking
- Excellent communication, able to explain complex AI platform topics to engineers and executives
- Structured, collaborative, cross-functional mindset
- Ability to produce reusable technical assets and runbooks
- Kubernetes (Docker/containerd, Helm, Operators)
- NVIDIA AI Enterprise components (NIM, NeMo, Triton, TensorRT, CUDA)
- GPU Operator and Kubernetes-based GPU lifecycle management
Reference: WJ-747_30137865