benchmark

3,501 ofertas de empleo de benchmark en México. Encuentra ofertas actualizadas diariamente de los principales portales de empleo.


  • Mexico City Engg Jornada completa

    MongoDB in Mexico City is seeking a Software Engineer to join the PM Benchmarking team. You will design and run performance benchmarks across cloud environments, build tooling, and analyze results to drive product decisions.The role emphasizes hands-on coding, ownership, and cross-functional collaboration within a hybrid work model. The ideal candidate has...

  • History PhD

    Hace 20 horas


    San Miguel Topilejo, México Mercor Jornada completa

    About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey. Position: Applied History & Political Science Benchmark Specialist Type: Contract Compensation:...


  • Ciudad de México Engg Jornada completa

    MongoDB in Mexico City is seeking a Software Engineer to join the PM Benchmarking team. You will design and run performance benchmarks across cloud environments, build tooling, and analyze results to drive product decisions. The role emphasizes hands-on coding, ownership, and cross-functional collaboration within a hybrid work model. The ideal candidate has...

  • History PhD

    Hace 20 horas


    San Antonio Tecómitl, México Mercor Jornada completa

    About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey. Position: Applied History & Political Science Benchmark Specialist Type: Contract Compensation:...

  • History PhD

    Hace 1 día


    distrito federal, México Mercor Jornada completa

    About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey. Position: Applied History & Political Science Benchmark Specialist Type: Contract Compensation: $44–$56/hour Location:...

  • History PhD

    Hace 2 días


    Ciudad de México Mercor Jornada completa

    About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey. Position: Applied History & Political Science Benchmark Specialist Type: Contract Compensation:...

  • History PhD

    Hace 2 días


    Mexico City Mercor Jornada completa

    About the job Mercorconnects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors includeBenchmark ,General Catalyst ,Peter Thiel ,Adam D'Angelo ,Larry Summers , andJack Dorsey . Position:Applied History & Political Science Benchmark Specialist Type:Contract Compensation:$44–$56/hour...


  • México Turing Jornada completa USD 23 Indefinido

    8-12 Weeks8+ yrs of relevant ExperienceComfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.Responsibilities: Validate task quality: Check that instructions, source materials, reference solutions, and evaluation criteria are consistent, with no hidden requirements or missing...


  • Ecatepec de Morelos, Estado de México Dicka Logistics Jornada completa

    DICKA Logistics busca un Analista de Compensaciones para integrarse al equipo, con foco en análisis de datos, nómina y estructuras salariales en un entorno nacional.Requiere Licenciatura y 1–2 años de experiencia; se valora dominio de Excel avanzado para tablas dinámicas y modelado de datos. Se ofrece oportunidad de desarrollo profesional en...


  • distrito federal, México Engg Jornada completa

    MongoDB in Mexico City is seeking a Software Engineer to join the PM Benchmarking team. You will design and run performance benchmarks across cloud environments, build tooling, and analyze results to drive product decisions.The role emphasizes hands-on coding, ownership, and cross-functional collaboration within a hybrid work model. The ideal candidate has...


  • mexico Lilt Jornada completa

    Lilt seeks experienced native-speaking software engineers to design, build, and validate multilingual benchmarks for evaluating large language models. This remote, freelance role focuses on high-signal tasks in your native language to test multilingual robustness without English translation crutches.You will handle task engineering, asset creation,...


  • mexico eDataBae Jornada completa

    eDataBae is seeking experienced AI Benchmark Quality Reviewers to support an advanced AI evaluation project focused on improving the quality and reliability of AI-generated solutions. You will review tasks, analyze model outputs, validate grading logic, and provide evidence-based feedback to enhance benchmark quality.The role requires 5+ years of relevant...


  • México eDataBae Jornada completa $391,000 - $725,000 Por obra

    eDataBae is seeking experienced AI Benchmark Quality Reviewers to support an advanced AI evaluation project focused on improving the quality and reliability of AI-generated solutions. You will review tasks, analyze model outputs, validate grading logic, and provide evidence-based feedback to enhance benchmark quality.The role requires 5+ years of relevant...


  • San Antonio Tecómitl, México Lilt Inc. Jornada completa

    About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal...


  • distrito federal, México Lilt Inc. Jornada completa

    About The OpportunityWe are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal...


  • San Miguel Topilejo, México Lilt Inc. Jornada completa

    About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal...


  • Ciudad de México Lilt Inc. Jornada completa

    About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal...


  • Mexico City Lilt Inc. Jornada completa

    About The OpportunityWe are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows....


  • Mexico (Remote), Mexico City LILT Trabajo remoto Jornada completa

    About The OpportunityWe are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal...


  • México Lilt Trabajo remoto Jornada completa $344,000 - $689,000 Por obra

    Lilt seeks experienced native-speaking software engineers to design, build, and validate multilingual benchmarks for evaluating large language models. This remote, freelance role focuses on high-signal tasks in your native language to test multilingual robustness without English translation crutches.You will handle task engineering, asset creation,...