ML Hardware Achitect
Job description
Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Tel Aviv, Israel; Haifa, Israel . Minimum qualifications:
• Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, a related field, or equivalent practical experience.
• 15 years of experience in computer architecture, ML accelerator design, or high-performance processor architecture.
• Experience leading architectural definition and authoring architecture specifications for silicon or compute IP blocks.
• Experience with performance modeling, workload profiling, and hardware-software co-design.
Preferred qualifications:
• Master's degree or PhD in Electrical Engineering, Computer Engineering, or Computer Science with an emphasis on computer architecture or ML hardware systems.
• 5 years of experience leading the architectural definition and microarchitecture of AI/ML accelerators from concept through production.
• Deep knowledge of modern deep learning workloads (Transformers, MoE, Diffusion, Generative AI inference) and their system bottlenecks (memory capacity, KV cache bandwidth, interconnect scaling).
• Strong understanding of high-performance memory subsystems (custom SRAM architectures, high-bandwidth memory hierarchies, caching schemes).
• Experience working with modern ML frameworks (PyTorch, JAX, TensorFlow) and ML compilers/runtimes (XLA, TVM, Triton).
About the job
In this role, you’ll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to drive cutting-edge TPU (Tensor Processing Unit) technology that powers Google's most demanding AI/ML applications. You’ll be part of a team that pushes boundaries, developing custom silicon solutions that power the future of Google's TPU. You'll contribute to the innovation behind products loved by millions worldwide, and leverage your design and verification expertise to verify complex digital designs, with a specific focus on TPU architecture and its integration within AI/ML-driven systems.
In this role, you will help shape the future of Google Cloud’s next-generation AI infrastructure, architecting high-performance Machine Learning silicon designed to power hyperscale AI inference. You will have an opportunity to drive accelerator technology that powers Generative AI models, large language models (LLMs), and emerging agentic workloads where throughput, latency, memory bandwidth, and energy efficiency are mission-critical.
You will be part of a silicon architecture team pushing the boundaries of custom computing. Leveraging your deep expertise in hardware-software co-design, machine learning algorithms, and computer architecture, you will define and optimize custom compute engines and memory hierarchies that accelerate the world's most advanced AI models across Google Cloud datacenters.
The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.
We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Responsibilities
• Lead the architectural definition, modeling, and specification of next-generation, high-performance ML compute IP and acceleration blocks for Cloud AI silicon.
• Own the ML IP architecture specification throughout the entire product lifecycle: concept exploration, cycle-accurate modeling, implementation, silicon bring-up, and production.
• Partner closely with leading AI research and algorithm teams (e.g., Google DeepMind, Gemini research teams) and software compiler teams (XLA, PyTorch) to explore architectural trade-offs and define hardware requirements for emerging model architectures.
• Drive comprehensive architecture studies, evaluating compute dataflows, numerical formats, sparsity, and specialized acceleration mechanisms such as key-value (KV) cache optimization.
• Drive performance, latency, power efficiency, and silicon area projections across model topologies and workload configurations.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .