evaluation

1,000 ofertas de empleo de evaluation en México. Encuentra ofertas actualizadas diariamente de los principales portales de empleo.

  • Research Scientist

    Hace 4 horas


    león, guanajuato, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    monterrey, nuevo león, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    zapopan, jalisco, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    tlalnepantla, estado de méxico Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    Ciudad de México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    juárez, chihuahua, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    guadalajara, jalisco, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    apodaca, nuevo león, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    culiacán, sinaloa, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    aguascalientes, aguascalientes, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    cuauhtémoc, sonora, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    iztacalco, distrito federal, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    hermosillo, sonora, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    san luis potosí, san luis potosí, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    el salto, jalisco, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    metepec, estado de méxico Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    venustiano carranza, distrito federal, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    xico, veracruz, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    celaya, guanajuato, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...

  • Research Scientist

    Hace 4 horas


    cuajimalpa, distrito federal, México Anyone Ai Jornada completa

    Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs — Human Data Division Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or...