evaluation

1,000 ofertas de empleo de evaluation en México. Encuentra ofertas actualizadas diariamente de los principales portales de empleo.

  • Research Scientist

    Hace 3 días


    Delegación Cuajimalpa de Morelos, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Colonia Nuevo Michoacán, Guanajuato, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Santiago Mexquititlán Barrio Primero, Querétaro, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    San Luis Potosí City, San Luis Potosí, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Pachuca, Hidalgo, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Álvaro Obregón, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Puebla City, Puebla, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Los Ángeles, Sonora, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Xico, Veracruz, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Guadalajara, Jalisco, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Tijuana, Baja California, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Naucalpan, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Iztapalapa, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Miguel Hidalgo, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Apodaca, Nuevo León, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Mexicali, Baja California, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Aguascalientes, Aguascalientes, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Toluca, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Cuauhtémoc, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Guadalupe, Nuevo León, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...