Perplexity · San Francisco · 🌐 Remote
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. The Answer Quality team ensures that our prompts, tools, search systems, datasets, and models work together to create the best possible experience for our users.As our product and agent capabilities evolve, we need evaluation systems that are fast, reliable, production-faithful, and actionable. In this role, you will build and improve the technical foundations that support Answer Quality across Perplexity. This includes our shared evaluation infrastructure and the platform used to replay and analyze agent traces. You will work closely with data scientists, engineers, and product teams to identify quality problems, measure their impact, and turn evaluation findings into product improvements.ResponsibilitiesBuild shared evaluation infrastructure that helps teams run reliable evals, analyze results, and make product and model decisionsDevelop the platform for replaying and analyzing agent traces to reproduce production behavior and diagnose failuresBuild and operate scalable systems for processing, storing, and monitoring interaction, trace, and evaluation dataPartner with data scientists, engineers, and product teams to turn answer-quality problems into evaluations, analyses, and product improvementsOperate in a small, high-impact team where your work directly shapes how Perplexity measures and improves Answer QualityQualifications4+ years of software, data, or machine learning engineering experience shipping and operating production systemsStrong proficiency in Python and SQL, with solid fundamentals in system design, data modeling, and distributed systemsExperience building big-data systems, including distributed compute, large-scale storage, and high-volume pipelinesDemonstrated ownership of ambiguous technical projects from initial design through production operationAbility to work effectively with data scientists, engineers, and product partnersPreferred QualificationsExperience building evaluation, experimentation, observability, or machine learning infrastructureFamiliarity with LLM and agent systems, including tool use, execution traces, replay, and simulationExperience building on top of large-scale data processing platforms such as Databricks, Snowflake, or ClickHouse
Want jobs like this matched to your resume, free? Create a free account — daily alerts, AI resume review, zero cost, forever.