Staff Software Engineer, AI and Infrastructure
Google · San Jose, US
Job description
Minimum qualifications:
- Bachelor's degree or equivalent practical experience.
- 8 years of experience programming in C++.
- 5 years of experience testing, and launching software products.
- 5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture.
- 3 years of experience with software design and architecture.
Preferred qualifications:
- Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
- 3 years of experience in a technical leadership role leading project teams and setting technical direction.
- 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.
- Experience with Linux Internals, Cluster Management, System Architecture, Virtualization, and Security.
About the job
Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren’t just building technology; you’re shaping the frontier of enterprise and driving the evolution of advanced models.
Our team develops Borglet which is Google’s “node agent”, responsible for managing the life cycle of all user processes that run on all our machines. The Borglet Infrastructure group in Borglet is a team of 30 SWEs (distributed between Sunnyvale and Warsaw) developing core pieces of Borglet (Software Architecture, Runtime, Machine Management, Storage) to deliver the container infrastructure that is scalable, extensible, efficient and secure. The team is focused on key large initiatives in MSCA like ML/AI Infrastructure (GPU/TPUs in cluster management system), Security, Capacity Fungibility, Warp Space (TI VMs).
The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.
We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $301000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities
- Design, implement, and analyze computer systems and their interactions with the kernel and hardware.
- Collaborate with partner teams as well as users across Google e.g., Borg team, ML teams, HW platform teams, SRE teams, Google's internal and Cloud users
- Solve ambiguous and high impact problems.
- Conduct strategic planning and tactical execution in complex projects. Ability to cross-coordinate across partner teams in Warsaw.
- Develop junior engineers on the team.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.
ML/AI Work links you to the employer's original posting — always verify the details there before applying.
More AI Security roles
View all →AI Security Specialist
CAA Club Group · Mississauga, CA
Senior Solutions Engineer: Public Sector & Defense AI
— · Baltimore, US
Forward Deployed Engineer - Clearance Required
LMI · Remote · Honolulu
AI Security Engineer
Barclays · Leeds, GB
Data Science Program Owner
BD · Remote · Carlow
Principal Agentic AI Security Engineer
AbbVie · Milwaukee, US