General information
Organisation
The French Alternative Energies and Atomic Energy Commission (CEA) is a key player in research, development and innovation in four main areas :
• defence and security,
• nuclear energy (fission and fusion),
• technological research for industry,
• fundamental research in the physical sciences and life sciences.
Drawing on its widely acknowledged expertise, and thanks to its 16000 technicians, engineers, researchers and staff, the CEA actively participates in collaborative projects with a large number of academic and industrial partners.
The CEA is established in ten centers spread throughout France
Reference
2026-41932
Description de l'unité
The French Alternative Energies and Atomic Energy Commission (CEA) is a leading research and innovation organization active in three major fields: energy, information and health technologies, and defense. Within the CEA, the Laboratory for Systems and Technology Integration (LIST), located in Saclay (Ile-de-France), drives technology transfer and innovation in embedded systems. The Embedded Artificial Intelligence Laboratory (LIAE) specializes in designing optimized AI solutions for embedded environments, with a focus on efficiency in surface, power consumption, and computing performance.
Position description
Category
Engineering science
Contract
Internship
Job title
KV Cache design for efficient transformer deployment in edge systems
Subject
The enthusiasm for neural network models like Transformers, dedicated to so-called generative AI tasks, remains undiminished. Thus, numerous models have been made available by the industrial in the market. Relying on the availability and exploitation of immense databases and costly deep learning, they achieved very high performances for content generation on GPU architectures. However, these large models are memory-intensive and very energy-intensive, making their use in the embedded domain very challenging. For LLM, the main bottleneck lies notably in massive access to memory. A so-called "KV Cache" module is essential to mitigate this problem and ensure an acceptable response generation time. It notably limits access to external memories by reusing data already calculated previously in the model stages. A large part of the work today focuses on its management.
Contract duration (months)
6
Job description
The LIAE is interested in studying and assessing the feasibility of implementing such a functionality. The main idea is to analyze and quantify the needs and limits for deployment on different architectural candidates (e.g. CPUs, GPUs, or SoCs...)
In this context, the objective of this internship is to:
- Provide a cross-disciplinary and well-supported state of the art of the needs of carefully selected LLMs and the various methods available for KV Cache described in the literature,
- Develop and validate a lightweight functional model in C++ within the laboratory's tool [11],
- Implement the mechanism in a selected LLM model.
Methods / Means
Python/Pytorch, développement en C++, Git
Applicant Profile
Engineering or Master’s 2 degree students
Skills: Initial experience with PyTorch and C++; implementation of inference; knowledge of hardware architectures would be an advantage
Documents to be provided: CV + Cover letter + Academic transcripts for the last 3 years
Position location
Site
Saclay
Job location
France, Ile-de-France
Location
Saclay
Candidate criteria
Prepared diploma
Bac+5 - Diplôme École d'ingénieurs
Recommended training
IA, embedded systems
PhD opportunity
Oui
Requester
Position start date
01/02/2027