Network Engineer
Job description
<div class="content-intro"><p><span style="font-family: arial, helvetica, sans-serif;">SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. </span><span style="font-family: arial, helvetica, sans-serif;">Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. </span><span style="font-family: arial, helvetica, sans-serif;">We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. </span><span style="font-family: arial, helvetica, sans-serif;">All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.</span></p></div><h3>ABOUT THE ROLE:</h3>
<p>SpacexAI is building and operating large-scale networks that underpin training and inference infrastructure, including high-performance / supercompute fabrics that connect GPU clusters, plus the core, edge, and datacenter networks that keep that infrastructure reachable and reliable.</p>
<p>We need a Network Engineer who is strong on fundamentals and comfortable owning production network design, deployment, and operations end to end. This is a hands-on engineering seat — not a NOC technician role and not a network-software (telemetry/ZTP platform) SWE role. You will design and build networks, qualify platforms, ship changes safely, and keep availability and performance high as we scale.</p>
<p>Travel to Memphis (and other build sites) may be required for capacity build-outs. You will participate in a team on-call rotation.</p>
<h3><strong>RESPONSIBILITIES:</strong></h3>
<ul>
<li>Design, deploy, and operate production datacenter and campus/core networks at scale</li>
<li>Own routing and switching configuration standards (BGP and at least one IGP such as OSPF or IS-IS), including change design, peer reviews, and execution</li>
<li>Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning</li>
<li> Build and improve monitoring, alerting, and operational documentation so issues are caught and fixed quickly</li>
<li>Troubleshoot Layer 2/Layer 3 incidents end to end — from link flaps and optics through routing and traffic engineering — and drive root cause and lasting fixes</li>
<li>Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil</li>
<li> Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows</li>
<li> Support high-performance / supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA-capable designs) as part of the broader network estate — deep specialist RoCE/NCCL ownership is a plus, not the bar for this seat</li>
</ul>
<h3><strong>BASIC QUALIFICATIONS:</strong></h3>
<ul>
<li> Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment</li>
<li> Solid hands-on experience with BGP and at least one interior routing protocol</li>
<li>Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics / high-speed Ethernet</li>
<li>Experience troubleshooting live production network incidents and participating in on-call</li>
<li>Strong written and verbal communication; clear change docs and incident notes</li>
</ul>
<h3><strong>PREFERRED SKILLS AND EXPERIENCE:</strong></h3>
<ul>
<li>Experience with modern datacenter vendors (e.g. Arista, Cisco, Juniper, Nvidia/Mellanox)</li>
<li>Familiarity with high-performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics) — useful context for our environment, not a hard filter</li>
<li> Network automation (Python, Ansible, Terraform, or similar) used in production</li>
<li>Experience with EVPN, leaf-spine, and large-scale Ethernet fabrics</li>
<li> Prior work supporting rapid datacenter or cluster capacity build-outs</li>
</ul>
<h3><strong>ADDITIONAL REQUIREMENTS:</strong></h3>
<ul>
<li>Willing to work onsite in Palo Alto</li>
</ul>
<h3>COMPENSATION AND BENEFITS:</h3>
<p>$150,000 - $250,000 USD</p>
<p>Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.</p><div class="content-conclusion"><p><em>SpaceXAI is an equal opportunity employer. For details on data processing, view our </em><em><a href="https://x.ai/legal/recruitment-privacy-notice" target="_blank">Recruitment Privacy Notice</a>.</em></p></div>