View job here

DATAGROUP and SMH Analytics have been working closely together in the field of Data Analytics for several years. As we continue to expand this successful partnership, we are now looking to strengthen our teams.

 

DATAGROUP is one of Germany’s leading IT service providers, driven by the expertise and commitment of our dedicated team. Our passion for technology and innovation enables us to provide our customers with reliable, high-quality IT services. At DATAGROUP, you can expect a fascinating and challenging role within an innovative company. An open corporate culture, individual training opportunities, and a wide range of professional development options provide a working environment that can be tailored to you and your life.

 

SMH Analytics GmbH is a consulting and training company operating at the intersection of Data Analytics, Artificial Intelligence, and Quality Management. Our clients range from German mid-sized companies to international industrial corporations. As the use of AI agents continues to grow across our client projects, we are expanding our Quality Assurance team – and we are looking for you.

Your Responsibilities:

  • You develop automated evaluation pipelines and test harnesses in Python (pytest, asyncio, pydantic) to reproducibly assess the behavior of GenAI agents.
  • You build golden datasets and task-based benchmarks, generate synthetic test cases, and work with our clients to define measurable acceptance criteria.
  • You evaluate not only final results, but the agent’s entire problem-solving process (trajectory testing), including correct tool selection, valid parameters, and efficient solution paths.
  • You design and calibrate LLM-as-a-Judge approaches against human evaluations and ensure the statistical reliability of the results.
  • You operationalize quality metrics such as Task Success Rate, Tool-Call Accuracy, Groundedness, and Robustness, including confidence intervals and sound sampling designs.
  • You set up observability and tracing solutions, for example with Langfuse, LangSmith, or OpenTelemetry, develop error taxonomies, and monitor quality in production environments.
  • You integrate evaluation suites as quality gates into CI/CD pipelines and conduct regression testing whenever prompts, tools, or models are changed.
  • You conduct red-team testing, including prompt injection and guardrail testing, and document testing procedures with regulatory requirements such as the EU AI Act in mind.
  • You prepare and communicate your findings in German for our clients’ specialist departments and management teams – clearly, precisely, and in a way that supports decision-making.

Our requirements:

  • Excellent Python skills: You write robust, production-ready code using pytest, asyncio, pandas, and pydantic. Python is part of your daily engineering toolkit, not just a scripting language.
  • Native-level German proficiency (C2): You communicate with our clients, write reports, and present results entirely in German.
  • Strong proficiency with agentic AI development tools such as Claude Code: You use these tools productively in your day-to-day development work, understand their strengths and limitations, and apply this hands-on experience directly to the design of your testing approaches.

Additional Qualifications That Will Make You Stand Out:

  • Experience with LLM SDKs such as Anthropic and OpenAI, as well as agent frameworks such as LangGraph
  • A solid foundation in statistics, including hypothesis testing, sampling design, and significance testing
  • Hands-on CI/CD experience with tools such as GitHub Actions or GitLab CI, as well as a basic understanding of IT security
  • Experience with KNIME or Power BI

What We Offer:

  • Remote-first: Work wherever you are most productive, complemented by regular in-person team meetings.
  • A full-time, permanent position in an area with enormous growth potential: quality assurance for AI agents is rapidly becoming business-critical for our clients.
  • Short decision-making paths and real responsibility: You work directly with the management team and help shape our methodology from the ground up.
  • A modern tooling environment: access to the latest AI tools, plus a dedicated budget for professional development and conferences.