Python Engineer, AI Coding Agent Evaluator
Anywhere
Workplace: RemoteContractUSD 100 - 200 hourlyFunction: Data Science & Machine LearningSkills: ["Engineering judgment","Opinionated feedback","Critical evaluation","Subjective rigor","Communication"]Evaluate the end-to-end quality of AI coding agent interactions (e.g., OpenAI Codex and Claude Code) in real-world scenarios. You’ll judge whether responses make sense, whether preambles and reasoning are useful, and whether outputs reflect strong engineering judgment and taste—not syntax correctness. Provide clear, opinionated feedback on what worked, what didn’t, and what feels misleading, helping define what “great” looks like with tools like Cursor.
Loading
Loading job details...
Preparing the role view and application actions.

