Senior GenAI Engineer

look4itRemote JobRemotecontractSenior
Active

Job description

This is a remote position. We are looking for a Senior AI Engineer to join a team building and operating a production-grade LLM system used by real users. This is not a proof-of-concept project. You will be responsible for taking LLM-powered features from idea through development and deployment to continuous improvement, with a strong focus on answer quality, performance, observability and cost efficiency. You will work with modern LLM technologies, including LangGraph, RAG, tool calling, Azure OpenAI, Gemini and Claude, and have a real impact on how AI-powered products are built and operated in production. Responsibilities: Build LLM-powered features end to end. Design and implement agentic flows, retrieval, and tool calling using LangGraph - then ship them as FastAPI services with streaming, persistence and proper tests. Own answer quality. Build evaluation datasets, regression suites and LLM-as-judge checks so we know whether a prompt or model change made things better before it reaches users. Get the right context to the model. Turn user questions into effective queries against our search platform, orchestrate multi-step research loops, and shape the context the model reasons over. When an answer is wrong, work out whether retrieval, the query or the prompt is at fault - and fix the right one. Debug production. Instrument flows with tracing (Langfuse), investigate bad answers from real traces, and manage latency, token and cost budgets - including routing across model sizes and families behind an AI gateway. Requirements Experience in building and operating backend services - APIs, async, testing Hands-on experience taking LLM features to production and keeping them running - not only prototypes Agent / orchestration frameworks - LangGraph ideally Practical RAG experience Experience debugging LLM systems in production - tracing, evaluation, cost and latency Experience running services in the cloud (we're on Azure) Strong problem-solving skills, analytical thinking, and technical decision-making Fluent in English, proactive communicator, and a collaborative team player Open-minded, creative, and motivated to push boundaries in AI and automation Nice to have: Azure OpenAI, AI Search, App Service Infrastructure-as-code (Bicep) Mentoring or tech-lead experience Tech stack: Python FastAPI LangGraph Azure OpenAI, Google Gemini, Anthropic Claude FAISS PostgreSQL Langfuse Azure App Insights Bicep Docker/Podman Benefits B2B contract 100% remote work Long-term engagement Work on a live LLM product Modern AI/LLM technology stack Flexible working environmentThis is a remote position. We are looking for a Senior AI Engineer to join a team building and operating a production-grade LLM system used by real users. This is not a proof-of-concept project. You will be responsible for taking LLM-powered features from idea through development and deployment to continuous improvement, with a strong focus on answer quality, performance, observability and cost efficiency. You will work with modern LLM technologies, including LangGraph, RAG, tool calling, Azure OpenAI, Gemini and Claude, and have a real impact on how AI-powered products are built and operated in production. Responsibilities: Build LLM-powered features end to end. Design and implement agentic flows, retrieval, and tool calling using LangGraph - then ship them as FastAPI services with streaming, persistence and proper tests. Own answer quality. Build evaluation datasets, regression suites and LLM-as-judge checks so we know whether a prompt or model change made things better before it reaches users. Get the right context to the model. Turn user questions into effective queries against our search platform, orchestrate multi-step research loops, and shape the context the model reasons over. When an answer is wrong, work out whether retrieval, the query or the prompt is at fault - and fix the right one. Debug production. Instrument flows with tracing (Langfuse), investigate bad answers from real traces, and manage latency, token and cost budgets - including routing across model sizes and families behind an AI gateway.

This is a remote position.

We are looking for a Senior AI Engineer to join a team building and operating a production-grade LLM system used by real users. This is not a proof-of-concept project. You will be responsible for taking LLM-powered features from idea through development and deployment to continuous improvement, with a strong focus on answer quality, performance, observability and cost efficiency. You will work with modern LLM technologies, including LangGraph, RAG, tool calling, Azure OpenAI, Gemini and Claude, and have a real impact on how AI-powered products are built and operated in production.Senior AI EngineerResponsibilities:Responsibilities:
  • Build LLM-powered features end to end. Design and implement agentic flows, retrieval, and tool calling using LangGraph - then ship them as FastAPI services with streaming, persistence and proper tests.
Build LLM-powered features end to end. Design and implement agentic flows, retrieval, and tool calling using LangGraph - then ship them as FastAPI services with streaming, persistence and proper tests.
  • Own answer quality. Build evaluation datasets, regression suites and LLM-as-judge checks so we know whether a prompt or model change made things better before it reaches users.
Own answer quality. Build evaluation datasets, regression suites and LLM-as-judge checks so we know whether a prompt or model change made things better before it reaches users.
  • Get the right context to the model. Turn user questions into effective queries against our search platform, orchestrate multi-step research loops, and shape the context the model reasons over. When an answer is wrong, work out whether retrieval, the query or the prompt is at fault - and fix the right one.
Get the right context to the model. Turn user questions into effective queries against our search platform, orchestrate multi-step research loops, and shape the context the model reasons over. When an answer is wrong, work out whether retrieval, the query or the prompt is at fault - and fix the right one.
  • Debug production. Instrument flows with tracing (Langfuse), investigate bad answers from real traces, and manage latency, token and cost budgets - including routing across model sizes and families behind an AI gateway.
Debug production. Instrument flows with tracing (Langfuse), investigate bad answers from real traces, and manage latency, token and cost budgets - including routing across model sizes and families behind an AI gateway.