Job description
Overview
In this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on. Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on. Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on. Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoE
In this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build and operate the AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AI CoEIn this role, you will work in a fast-paced, agile environment with a diverse team that has a true passion for technology, transformation, and outcomes. You will help build andoperatethe AWS-based platform services that power our multi-agent AI workflows for the enterprise, as part of the AICoEThis is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on.
This is a senior, hands-on engineering role. You will implement and operate the core services of the AI/Agentic platform; building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AI Platform Lead into reliable, scalable systems that other engineering and agent teams build on.This is a senior, hands-on engineering role. You will implement andoperatethe core services of the AI/Agentic platform;building backend, orchestration, and shared services (RAG, memory, human-in-the-loop), deploying them as infrastructure-as-code, and hardening them for production. You will translate the architecture and standards set by the AIPlatform Leadinto reliable, scalable systems that other engineering and agent teams build on.Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomes
Here at Waters, we look to our team members to be versatile and enthusiastic about tackling new problems, display strong ownership, and remain focused on business outcomesHere at Waters, we look to our team members to be versatile and enthusiastic about tacklingnew problems, display strong ownership, and remain focused on business outcomesResponsibilities
Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoptionBuild, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoptionBuild, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption- Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination Package & deploy platform services via IaC onto Cloud Ops-managed foundations Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption
- Build, deploy, and operate AI/Agentic platform services in AWS and the orchestration systems that power multi-agent AI workflows, including agent lifecycle, routing, and coordination
- Package & deploy platform services via IaC onto Cloud Ops-managed foundations
- Build and operate the shared platform capabilities that agent teams consume, RAG, memory, and human-in-the-loop services
- Build & maintain CI/CD and LLMOps/AgentOps pipelines so AI services stay reliable and scalable in production; apply the CoE standards and contribute improvements back
- Build and operate evaluation and quality tooling for AI services; golden datasets, LLM-as-judge and groundedness/hallucination checks, and prompt/agent regression
- Own the reliability of shared platform services: monitoring, logging, and observability for model and agent performance and health; and responding to and resolving production issues
- Apply the platform’s governance and security controls in the services you build, cost/spend guardrails, connector permission scope, and prompt-injection and agent-attack mitigation
- Design reusable, well-documented service abstractions, templates, and SDKs that application and agent teams build on, and partner with those teams to bring new AI capabilities into production and continually improve the experience of building on the platform
- Provide constructive code reviews and mentor junior and mid-level engineers to ensure the engineering standards across the platform team
- Document platform services, interfaces, and patterns so consuming teams can self-serve, and support knowledge-sharing across the AI CoE
- Stay abreast of AI industry trends and proactively identify opportunities for improvement and adoption
Qualifications
Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practicesBachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practicesBachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices- Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience. 5+ years of industry experience in software engineering or ML Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents). Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack Working knowledge of vector databases, memory systems, and human-in-the-loop workflows. Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices
- Bachelor or Master degree in CS, AI/ML, Data Science, or equivalent practical experience.
- 5+ years of industry experience in software engineering or ML
- Working understanding of agent frameworks and architectures (e.g., LangChain, LangGraph, or CrewAI), sufficient to design, build, and operate the platform services that support them, including RAG and memory
- Strong programming skills in one or more languages (e.g., Python, Java, Go, Node.js, or TypeScript), with solid experience building backend services and distributed systems in production, including microservices and event-driven architectures, and designing and versioning REST or GraphQL API contracts
- Solid experience using managed cloud services (AWS preferred: Bedrock/AgentCore, Lambda, API Gateway, eventing; GCP or Azure acceptable) to build backend and orchestration services
- Hands-on experience operating an agent runtime such as AWS Bedrock AgentCore (or an equivalent) is a strong plus
- Hands-on experience building CI/CD pipelines and applying LLMOps/AgentOps practices for deploying, scaling, and monitoring LLM-powered services
- Knowledge of LLM/agent evaluation and quality measurement (e.g., golden sets, LLM-as-judge, groundedness/hallucination metrics, regression testing of prompts and agents).
- Experience with monitoring/logging stacks such as Datadog, Grafana, or the ELK stack
- Working knowledge of vector databases, memory systems, and human-in-the-loop workflows.
- Strong collaboration skills across platform engineering and product teams, with clear communication and a bias toward shared standards and reusable solutions
- Curious mindset with a strong desire to stay ahead of AI/ML advancements and enterprise best practices