Neural Video Compression using autoencoders
Imagine describing a complex image using only 10 keywords, after that, someone else try to redraw the image leveraging the 10 keywords.
|
Computer Vision and Deep Learning Engineer — ESIR, Université de Rennes.
I build intelligent systems that understand, reconstruct, and analyze visual data. My work focuses on image processing, computer vision, information theory, and deep learning — with strong interests in visual reconstruction, inverse problems, and generative models.
What drives me: bridging theory and practice by combining mathematical modeling, research-driven experimentation, and scalable high-performance implementation.
Currently seeking a PhD or full-time position in computer vision, image processing, and machine learning.
Outside of work: basketball, football, and following the latest in AI research.
My work sits at the intersection of mathematics, signal processing, and deep learning. Here are the areas I think about most deeply.
Visual understanding, recognition, and reconstruction — from classical geometry to modern deep learning pipelines.
Exploring score-based models, diffusion processes, and VAEs for image synthesis and inverse problem solving.
Rate-distortion theory, entropy coding, and the mathematical foundations of image and video compression.
Neural video compression, codec design (JPEG, H.264, HEVC), and real-time processing pipelines.
Learning-based and geometry-based methods for recovering 3D structure from 2D observations.
Denoising, deblurring, super-resolution — using optimization and deep priors to recover signals from degraded observations.
CUDA kernels, parallel algorithms, and high-performance pipelines for compute-intensive visual workloads.
Segmentation, classification, and computer-aided diagnosis applied to clinical imaging datasets.
Multimodal biomass estimation from aerial imagery combining DINOv2 and tabular metadata.
ResNet18 model for computer-aided diagnosis on Chest X-Ray images — 84.93% accuracy.
U-Net architecture for pixel-level segmentation on medical imaging datasets.
Pipeline combining geometry-based and learning-based methods to recover 3D structure from 2D images.
MLP approximating multi-body gravitational potential in N dimensions.
Autoencoder architecture for image denoising and signal reconstruction.
LSTM model for many-to-one prediction on a simulated sine wave.
Reuters newswire topic classification across 46 categories using GRU.
CNN trained from scratch on CIFAR-10 with augmentation and TensorBoard — 70.02% accuracy.
Rendering pipeline for 3D models with OpenGL, Blender, and Assimp.
3D shape reconstruction from GCode with geometric analysis and comparison.
Thoughts on computer vision, deep learning, and research — written to share what I learn. Also published on Medium.
Imagine describing a complex image using only 10 keywords, after that, someone else try to redraw the image leveraging the 10 keywords.
I'm actively looking for a PhD position or full-time role in computer vision and ML. Let's talk.