You will work directly with business teams to understand the processes, decisions, systems, data, exceptions, and risks involved in each use case.
You will design the architecture of agent-based systems, selecting the most appropriate models, patterns, and components for each problem.
You will build applications powered by language models using context engineering, tool calling, structured outputs, RAG, and multi-step workflows.
You will experiment with models, instructions, and architectures to optimize the balance between quality, latency, cost, and reliability.
You will connect agents to internal and external tools through APIs, MCP servers, databases, and SaaS platforms.
You will rapidly create prototypes and evolve them into solutions used with real-world data, systems, and use cases.
You will design evaluations, test cases, and regression tests, and instrument agents to monitor their traces, errors, latency, and costs.
You will implement reliability and control mechanisms such as permissions, validations, retries, fallbacks, action limits, and escalation to a human.
You will deploy and operate solutions using testing, version control, CI/CD, and sound engineering practices.
You will turn the lessons learned from each deployment into reusable components, evaluations, and patterns for future Tuio agents.
You will become embedded in one of Tuio’s priority domains to understand its processes, systems, and key technical challenges. You will take responsibility for one or two use cases and bring at least one of them from prototype to a pilot with real users, defining its evaluations, observability, and next steps for scaling.
Your day-to-day work will combine close collaboration with the business, end-to-end technical ownership, and a way of building based on experimenting quickly, measuring rigorously, and strengthening controls whenever the level of risk requires it.
You will be involved from the discovery stage and work directly with the people who understand and operate each process. Depending on the project, you will collaborate with AI Managers and colleagues from Business, AI, Data, and Engineering to define and evolve the solution.
You will take end-to-end technical ownership, including AI architecture, code, integrations, evaluations, deployment, observability, and reliability. In some projects, you will also play a significant role in discovery and functional design.
You will build small initial versions to validate hypotheses with real users, measure outcomes, and iterate quickly. We will consciously accept technical debt when it helps us learn, but we will strengthen the architecture, testing, and security before scaling.
You will treat agents as production systems: you will version instructions, tools, and evaluations; analyze traces and failure modes; measure quality, latency, and cost; and test every change against representative cases.
You will adapt the speed of development and level of control to the degree of risk. In low-risk cases, you will experiment with a high level of autonomy. In more sensitive processes, you will introduce permissions, traceability, validations, action limits, data isolation, and human oversight.
You will turn the lessons learned from each project into reusable components, evaluations, and patterns, helping make every new agent faster and more reliable to build than the previous one.
The goal will not simply be to launch a demo, but to demonstrate that an agent can improve a real-world operation in a measurable, reliable, and responsible way.
