Lead Engineer – AI Platform

Waters CorporationJob LocationsSenior
ActiveVerified 8m ago

Job description

Overview

In this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on. Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on. Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on. Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomes

In this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE

In this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoEIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build andoperatethe AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AICoE

This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on.

This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on.This is a senior, hands-on engineering role. You will implement andoperatethe core services of the AI/Agentic platform;building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AIPlatform Leadinto reliable, scalable systems that other engineering and agent teams build on.

Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomes

Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesHere at Waters, we look to our team members to be versatile and enthusiastic about tacklingnew problems, display strong ownership, and remain focused on business outcomes

Responsibilities

Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoptionBuild, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoptionBuild, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption
  • Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption
  • Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination
Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordinationBuild, deploy, andoperateAI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination
  • Package & deploy platform services via IaC onto Cloud Ops-managed foundations
Package & deploy platform services via IaC onto Cloud Ops-managed foundationsPackage&deploy platform services viaIaConto Cloud Ops-managed foundations
  • Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services
Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop servicesBuild andoperatethe shared platform capabilities that agent teams consume,RAG, memory, and human-in-the-loop services
  • Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back
Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements backBuild&maintainCI/CD andLLMOps/AgentOpspipelines so AI services stay reliable and scalable in production; apply theCoEstandards and contribute improvements back
  • Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression
Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regressionBuild andoperateevaluation and quality tooling for AI services;golden datasets, LLM-as-judge andgroundedness/hallucination checks, and prompt/agent regression
  • Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues
Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issuesOwn the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues
  • Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation
Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigationApply the platform’s governance and security controls in the services you build,cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation
  • Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform
Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platformDesign reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform
  • Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team
Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform teamProvideconstructive code reviews and mentor junior and mid-level engineers toensure theengineering standards across the platform team
  • Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE
Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoEDocument platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AICoE
  • Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption
Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoptionStay abreast of AI industry trends and proactivelyidentifyopportunities for improvement and adoption

Qualifications

Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practicesBachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practicesBachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices
  • Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices
  • Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience.
Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience.Bachelor orMasterdegree inCS, AI/ML, Data Science, or equivalent practical experience.
  • 5+ years of industry experience in software engineering or ML
5+ years of industry experience in software engineering or ML5+ years of industry experience in software engineering orML
  • Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory
Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memoryWorking understanding of agent frameworks and architectures (e.g.,LangChain,LangGraph, orCrewAI),sufficient to design, build, andoperatethe platform services that support them, including RAG and memory
  • Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts
Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contractsStrong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-drivenarchitectures, and designing and versioning REST orGraphQLAPI contracts
  • Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services
Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration servicesSolid experienceusing managedcloudservices (AWS preferred:Bedrock/AgentCore, Lambda, API Gateway,eventing; GCP or Azure acceptable) to build backend and orchestration services
  • Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus
Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plusHands-on experienceoperatingan agent runtime such as AWS BedrockAgentCore(or an equivalent) is a strong plus
  • Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services
Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered servicesHands-on experience building CI/CD pipelines and applyingLLMOps/AgentOpspractices for deploying, scaling, and monitoring LLM-powered services
  • Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents).
Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents).Knowledge ofLLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge,groundedness/hallucination metrics, regression testing of prompts and agents).
  • Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack
Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stackExperience with monitoring/logging stacks such asDatadog,Grafana, or the ELK stack
  • Working knowledge of vector databases, memory systems, and human-in-the-loop workflows.
Working knowledge of vector databases, memory systems, and human-in-the-loop workflows.Working knowledge of vector databases, memory systems, and human-in-the-loop workflows.
  • Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions
Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutionsStrong collaboration skills across platform engineering and product teams, with clear communication and a biastoward shared standards and reusable solutions
  • Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices
Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practicesCurious mindset witha strong desireto stay ahead of AI/ML advancements and enterprise best practices