Stage Human-Object Interaction benchmarking-Saclay-H/F

Détail de l'offre

Informations générales

Entité de rattachement

Le CEA est un acteur majeur de la recherche, au service des citoyens, de l'économie et de l'Etat.

Il apporte des solutions concrètes à leurs besoins dans quatre domaines principaux : transition énergétique, transition numérique, technologies pour la médecine du futur, défense et sécurité sur un socle de recherche fondamentale. Le CEA s'engage depuis plus de 75 ans au service de la souveraineté scientifique, technologique et industrielle de la France et de l'Europe pour un présent et un avenir mieux maîtrisés et plus sûrs.

Implanté au cœur des territoires équipés de très grandes infrastructures de recherche, le CEA dispose d'un large éventail de partenaires académiques et industriels en France, en Europe et à l'international.

Les 20 000 collaboratrices et collaborateurs du CEA partagent trois valeurs fondamentales :

• La conscience des responsabilités
• La coopération
• La curiosité
  

Référence

2026-41776  

Description de l'unité

Based in Saclay (Essonne), the LIST is one of the two institutes of CEA Tech, the Technological Research Division of the CEA. Dedicated to intelligent digital systems, its mission is to carry out technological developments of excellence on behalf of industrial partners, in order to create value.
Within the LIST, the Laboratory of Vision and Learning for Scene Analysis (LVA) conducts its research in the field of computer vision and artificial intelligence for the perception of intelligent and autonomous systems. The laboratory's research themes include visual recognition, behavior and activity analysis, large-scale automatic annotation, and perception and decision models.

Description du poste

Domaine

Mathématiques, information  scientifique, logiciel

Contrat

Stage

Intitulé de l'offre

Stage Human-Object Interaction benchmarking-Saclay-H/F

Sujet de stage

The interpretation of human interactions in images or videos has significantly improved with the emergence of Large Language Models (LLMs) and Vision-Language Models (VLMs). However, these large models, whether used directly or distilled into specialized models, still have significant limitations, particularly in accurately attributing interactions to the correct person in dense scenes and discriminating actions in the presence of objects.

Evaluation protocols and databases for this task do not always accurately reflect the true capabilities of the methods due to issues such as annotation imprecision or overly rigid semantic metrics. This internship tackles this problem.

Durée du contrat (en mois)

6 mois

Description de l'offre

Context

The interpretation of human interactions in images or videos has significantly improved with the emergence of Large Language Models (LLMs) and Vision-Language Models (VLMs). However, these large models, whether used directly or distilled into specialized models, still have significant limitations, particularly in accurately attributing interactions to the correct person in dense scenes and discriminating actions in the presence of objects.
 
Evaluation protocols and databases for this task do not always accurately reflect the true capabilities of the methods due to issues such as annotation imprecision or overly rigid semantic metrics.

What do we expect from you?

To address these problems, the internship will focus on the following objectives:
- Conduct a state-of-the-art review of existing databases and analyze their biases (e.g., precision of detection boxes).
- Propose a semi-automatic pipeline for correcting these biases.
- Identify the biases and gaps in the metrics commonly used in the state-of-the-art.
- Propose a new benchmark, addressing various application domains.
- Evaluate the main state-of-the-art approaches on this benchmark.
- Write a publication about this benchmark.

 

#Cea List

Moyens / Méthodes / Logiciels

AI, Deep Neural Network, Computer Vision, Human behavior analysis

Profil du candidat

Profile

- Students in their 4th or 5th year of studies (M1, M2 or gap year) 
- Computer vision skills
- Machine learning skills (deep learning, perception models, generative AI…) 
- Python proficiency in a deep learning framework (especially TensorFlow or PyTorch)

Localisation du poste

Site

Saclay

Localisation du poste

France, Ile-de-France, Essonne (91)

Ville

Saclay

Critères candidat

Diplôme préparé

Bac+5 - Master 2

Formation recommandée

AI, Deep Learning, Computer Vision

Possibilité de poursuite en thèse

Oui

Demandeur

Disponibilité du poste

01/02/2027