2d ago

avatar

Leapsome

Staff Product Engineer (d/f/m)

€110K - €165K

Berlin, BE, Germany

Senior (10+ years)

SaaS

Growing (201–500)

[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object]

Questions about the Staff Product Engineer (d/f/m) role at Leapsome

How do you balance AI autonomy with human-in-the-loop safety in production?

To balance AI autonomy with human-in-the-loop safety at Leapsome, I prioritize a "trust but verify" architecture. I design systems where AI agents handle complex, routine orchestration—like proactive nudges or workflow triggers—while integrating explicit checkpoints for sensitive actions. By implementing strict schemas for tool calling and dedicated AI evaluation suites (LLM-as-a-judge), I ensure decisions align with business logic before execution. For high-stakes operations, I build "human-in-the-loop" fallbacks, requiring manual confirmation for data-sensitive tasks. This approach treats AI as a powerful force multiplier while maintaining the engineer’s accountability, ensuring our proactive workflows provide genuine ROI without compromising the security, reliability, or data integrity of our platform.

What key metrics define success for AI-driven features in this industry?

Success for AI-driven features at Leapsome is measured through both technical performance and tangible business impact. Key metrics include agentic performance outcomes, such as the successful execution of complex workflows and the accuracy of autonomous task completion within the HR lifecycle. To ensure high-quality, safe interaction, teams track "human-in-the-loop" intervention rates and AI evaluation scores (LLM-as-a-judge). Furthermore, success is defined by business ROI, measured through workforce insights, the efficacy of proactive triggers, and user engagement with AI-generated nudges. Ultimately, these metrics must prove that the AI tools reliably scale operations, reduce manual HR processes, and measurably improve employee development, engagement, and performance outcomes across the client platform.

How are you evolving RAG architectures to handle complex data context?

To evolve RAG architectures for complex data contexts, we move beyond simple retrieval to stateful, agentic orchestration. We leverage LangGraph to build multi-agent workflows capable of navigating interconnected HR data across our platform. By implementing advanced RAG pipelines, we ensure high-precision context retrieval through robust embedding strategies and vector databases.

We prioritize data integrity and security by architecting strict schemas for tool calling and incorporating "human-in-the-loop" safety fallbacks. To ensure genuine business ROI, we integrate LLM-as-a-judge frameworks (like LangSmith) to evaluate agent performance iteratively. This allows us to deliver proactive, intelligent nudges that synthesize complex employee information while maintaining the high standards required for enterprise HR systems.

How will this role shape your transition to proactive HR workflows?

As a Staff Product Engineer, you will transition the Leapsome platform from reactive tools to proactive workflows by architecting intelligent, autonomous systems. You will serve as the technical anchor, designing foundational context layers and robust security architectures that allow the platform to anticipate user needs. By integrating LLM-driven agents, RAG pipelines, and sophisticated tool-calling schemas, you will enable the platform to orchestrate complex HR processes autonomously. You will leverage agent performance analytics and workforce insights to refine these triggers, ensuring that "nudges" provide genuine, data-driven value. Ultimately, your work establishes the standard for human-AI collaboration, turning the platform into a proactive partner that empowers organizations to achieve their goals more efficiently.

How does the engineering culture support rapid AI-driven product iteration?

Leapsome fosters rapid AI-driven iteration by embedding AI into the core development lifecycle rather than treating it as an auxiliary tool. The culture prioritizes "pragmatic excellence," where engineers act as "owners without egos," taking full accountability for AI-generated code while using tools like Cursor, CodeRabbit, and Claude Code as force multipliers.

The team maintains a technical edge by pioneering dedicated AI evaluation suites (evals), strict schemas for tool calling, and "human-in-the-loop" safety fallbacks. By aligning product, design, and engineering, the culture supports high-speed prototyping and data-driven refinement. This allows engineers to rapidly deploy complex multi-agent workflows and RAG pipelines while maintaining enterprise-grade security and long-term architectural scalability.