Field Solutions Architect IV, AI Infrastructure, Google Cloud
Job description
In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees. Benefits for this role include:
• Health, dental, vision, life, disability insurance
• Retirement Benefits: 401(k) with company match
• Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
• Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
• Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
• Baby Bonding Leave: 18 weeks
• Holidays: 13 paid days per year
Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Sunnyvale, CA, USA; Atlanta, GA, USA; Chicago, IL, USA; Kirkland, WA, USA; Austin, TX, USA; New York, NY, USA . Minimum qualifications:
• Bachelor's degree in Science, Technology, Engineering, Mathematics, or equivalent practical experience.
• 10 years of experience with data center technologies, solutions and their deployment..
• Experience with data center networking, such as routing, switching, DNS, and IP addressing; compute and storage concepts; and operations including power, cooling, and procedures.
• Experience with AI model training and inference, deploying OSS frameworks such as PyTorch, Jax, etc., and optimizing performance versus costs.
• Experience with scripting and coding in Python, YAML, and Kubernetes.
Preferred qualifications:
• Experience training and fine tuning large models with accelerators.
• Experience with performance profiling and benchmarking tools (i.e., MLPerf, PyTorch profiler, Tensorboard).
• Experience deploying compute, network and storage solutions in data centers.
• Experience deploying TPU clusters.
About the job
As an AI Infrastructure Field Solution Architect, your experience and thought leadership will support Google Cloud commercial teams to incubate, pilot, and roll out industry-leading TPU and GPU hardware at market innovators, large enterprises, and early-stage startups. You will empower accounts to innovate faster with systems and software using our latest technical offerings.
You will guide technical discussions on network topologies, fabric architecture, and compute or storage integration, while supporting server and cluster provisioning. This includes on-site visits to client hosting facilities during physical implementation.
You will identify and assess opportunities, evaluating cost-to-performance tradeoffs, evaluating baseline models, developing migration paths, and integrating specialized silicon into overarching technical roadmaps. Along the way, you will partner closely with internal product and engineering units to remove roadblocks and shape future platform capabilities.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $233000 - $324000 (USD) + 25% bonus target + equity + benefits
Learn more about benefits at Google .
Responsibilities
• Act as a trusted advisor to customers, helping them understand and incorporate AI accelerators into their overall business strategy by designing training and inferencing platforms, and incorporating customers' existing data center technologies.
• Demonstrate how Google Cloud is differentiated, highlighting the power of accelerators by working with customers on Proofs of Concept, demonstrating features, optimizing model performance, profiling, and benchmarking.
• Assist in production deployment of Google Cloud accelerators, troubleshoot integration issues with existing compute, networking, and storage solutions, and enable customers to operate their training and inferencing clusters successfully.
• Build repeatable assets to enable other customers and internal teams.
• Collaborate with the AI Infrastructure Dedicated Engineering Team in Google Compute Engine, and scale ML techniques and solutions broadly. Influence Google Cloud strategy at the intersection of infrastructure and AI by advocating for enterprise customer requirements.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .