Сравним вакансию с вашим действующим резюме и составим письмо под эту роль.
Описание вакансии
Anthropic is an AI safety and research company building reliable, interpretable AI systems like the Claude family of large language models.
About the role
Prompts specify what Claude does. Evals measure whether it did. The Product Prompt and Eval Design team handles both for product design: the system prompt greeting new users, tool descriptions deciding whether Claude searches, instructions maintaining quality, and evals testing all of it. The role focuses on building evals that check prompts, the harness running them, and tools enabling designers to do this work independently. It involves close collaboration with surface owners, engineers, and prompt engineering teams during model releases.
Key responsibilities:
— Write and revise prompts behind Claude's tools, features, and behaviors on product surfaces; test, fix, ship, and confirm prompts.
— Build graders to validate prompt fixes and rerun on new models; automate evals from designers' rubrics and analyze transcripts.
— Develop visual, low-code eval tools for designers without engineering help: assemble comparison sets, convert rubrics into graders, compare prompt variants, and present results.
— Observe designers using tools and simplify them.
— Support model releases by testing surfaces, writing prompt fixes and migrations, and creating prompts for new features.
— Establish and scale the eval harness to test 50-100 tools with reliable settings, maintain eval stability across models, and diagnose regressions.
— Package issues unfixable by prompting for training, including graders, human-feedback questions, and good/bad pairs.
Minimum qualifications:
— Production-quality Python.
— Experience building and maintaining evaluation pipelines for LLM products: graders, rubrics, comparison sets, regression suites, and their orchestration.
— Experience creating internal tools with user-friendly interfaces for non-coders.
— Experience setting up test harnesses, sandboxing tool calls, and stabilizing settings for comparable runs.
— Experience shipping prompts or working closely with prompt shippers, understanding prompt failures across models.
— Ability to read transcripts beyond scores.
Preferred qualifications:
— Experience within a model-launch cycle.
— A/B testing experience linking offline evals to online outcomes.
— Front-end or notebook-to-app experience and insights on eval result legibility.
— Experience converting product rubrics into training signals: graders, human-feedback questions, or preference pairs.
— Concern for Claude's behavior for users, beyond metric changes.
Logistics:
— Minimum education: Bachelor’s degree or equivalent.
— Location-based hybrid policy: at least 25% office presence expected.
— Visa sponsorship available with legal support.
Anthropic values diversity and encourages applications from underrepresented groups. The company emphasizes AI safety and ethical implications, fostering a collaborative environment focused on impactful AI research.
Compensation:
— Annual salary range: $305,000 - $385,000 USD.
Work location:
— Hybrid in San Francisco or New York offices.
Контакты работодателя доступны по кнопке «Откликнуться» после входа.