Technical Site Reliability Engineer
Anduril Industries
As founding Site Reliability Engineer in the Advanced Capabilities team, you’ll design, build, and operate the simulation infrastructure that enables massive-scale autonomous-system wargaming. You’ll own the hardware-to-software stack, ensure reliable post-release validation, and prevent regressions that could impact real-time operations. You’ll work closely with development teams to keep the Simulation Center running, with a focus on stability, performance, and rapid incident response. This role blends hands-on engineering with cross-team collaboration in a cutting-edge defense tech environment.
Pay / Benefits- equity grants
- comprehensive benefits package
- health coverage and recovery support
- Maintain the simulation software stack across tools and environments
- Own and optimize compute, networking, storage, and configuration for the simulation
- Build and extend a post-release regression and smoke-test suite
- Diagnose, root-cause, and eliminate failure modes with guardrails and monitoring
- Partner with development teams to assess reliability risks before releases
- Monitor system health, triage issues, and escalate with context
- Document runbooks, issues, and configurations for operational knowledge
- Support release validation and environment configuration through automation and tooling
- Proficiency in Python for automation and test development
- Working knowledge of C++ for reading, debugging, and tracing simulation code
- Solid general networking fundamentals (TCP/IP, UDP, multicast, DNS, routing, firewalls) and distributed systems troubleshooting
- Experience with project management, issue tracking, bug triage, and cross-team coordination
- Demonstrated experience maintaining production or production-adjacent systems under time pressure
- Strong written and verbal communication; ability to document and escalate clearly
- Eligibility to pass security and background checks for sensitive information systems
- strong written and verbal communication
- ability to escalate and explain root causes
- proactive problem-solving
- Python automation
- C++ debugging and tracing
- networking fundamentals across distributed systems
Reference: WJ-747_30155758