evaluation

1,000 ofertas de empleo de evaluation en México. Encuentra ofertas actualizadas diariamente de los principales portales de empleo.

  • Research Scientist

    Hace 3 días


    General Escobedo, Nuevo León, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Zapopan, Jalisco, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Hermosillo, Sonora, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Azcapotzalco, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    San Pedro, Coahuila, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Cuautitlán Izcalli, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Naucalpan, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Lerdo, Durango, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    San Pedro Garza García (Jesús María), Nuevo León, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Tijuana, Baja California, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Benito Juarez, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Tlalnepantla, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Santiago Mexquititlán Barrio Primero, Querétaro, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    San Luis Potosí City, San Luis Potosí, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Venustiano Carranza, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Cancún, Quintana Roo, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Pachuca, Hidalgo, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Álvaro Obregón, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Ciudad Juárez, Chihuahua, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 3 días


    Iztapalapa, Mexico City Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...