9d ago

avatar

Snowflake

Senior Software Engineer - Cortex AI - FDE

$200K - $270K

Menlo Park, CA

Senior (10+ years)

SaaS

Enterprise (1000+)

[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object]

Questions about the Senior Software Engineer - Cortex AI - FDE role at Snowflake

What key technical metrics define success for an AI orchestration engine?

Success for an AI orchestration engine at Snowflake is defined by a balance of performance, reliability, and observability. Key metrics include latency, specifically focusing on the time-to-first-token and total execution speed for tool-use workflows. Reliability is measured through system uptime, error rates during high-throughput execution, and state management consistency. Operational efficiency is tracked via cost-per-request, prompt caching hit rates, and optimized token usage. Finally, system quality is assessed using "Evals" data: success rates of agentic workflows, accuracy of RAG retrieval-augmented generation pipelines, and the ability to diagnose issues rapidly through cross-layer tracing. Ultimately, these metrics ensure the infrastructure remains scalable, secure, and performant for enterprise-grade agentic applications.

How are industry standards for LLM evaluation evolving for production agents?

Industry standards for LLM evaluation are shifting from static, manual testing toward automated, continuous infrastructure. For production agents, "evals" now rely on golden set simulations and large-scale, iterative pipelines rather than anecdotal verification. Organizations are moving toward "hillclimbing" experiments where systemic quality metrics are tracked alongside latency and cost.

Furthermore, standards are evolving to integrate cross-layer debugging, where observability tools trace agent logic through retrieval and execution stages to root-cause failures. The industry now treats evaluation as a core engineering discipline, utilizing automated pipelines to quantify accuracy, guardrail adherence, and cost-efficiency in real-time, ensuring that agentic workflows remain hardened, reliable, and scalable in complex, multi-tenant enterprise environments.

What are the biggest scalability challenges in RAG systems today?

The biggest scalability challenges in RAG systems center on managing high-throughput context retrieval and maintaining performance as data volume grows. Key hurdles include designing efficient search indexing and ranking mechanisms that handle massive datasets without latency degradation. Furthermore, balancing real-time query processing with semantic caching requires sophisticated infrastructure. As systems scale, ensuring consistency in distributed state management while integrating complex vector databases becomes critical. Additionally, cost-effectively optimizing prompt caching and model routing is essential for enterprise-grade deployment. Ultimately, developers must build robust observability and automated metadata extraction pipelines to ensure that AI workflows remain reliable, secure, and performant while operating across multi-tenant environments at an industrial scale.

How does the Cortex team balance new feature velocity with platform stability?

The Cortex team balances rapid innovation with platform stability by building hardened, production-grade infrastructure that treats AI as a core architectural component rather than an experimental add-on. They achieve this velocity through the "Evals Engine," which enables massive-scale golden set simulations and automated error analysis to validate workflows before deployment.

Stability is maintained by treating LLM features as multi-tenant microservices with rigorous observability and guardrails. By focusing on low-latency orchestration, prompt caching, and infrastructure-led cost optimizations, the team ensures that new agentic capabilities scale reliably. Ultimately, their approach relies on cross-layer debugging and a product-focused mindset to distinguish between unique customer issues and platform gaps, ensuring long-term systemic reliability.

How does the team prioritize between bespoke customer needs and platform gaps?

The team prioritizes between bespoke customer needs and platform gaps by exercising "product instinct," a core competency for this role. Instead of reacting to every individual request, engineers must judge whether a specific customer's problem is truly unique or if it represents a broader structural deficiency in the platform. This involves critical decision-making to determine if a fix provides repeatable value for all users or merely solves a one-off issue. By distinguishing between customized work and platform-level requirements, the team ensures they build scalable, reusable infrastructure rather than fragmented, hard-to-maintain solutions, ultimately accelerating the evolution of their agentic AI ecosystem for the entire enterprise.