I'm an AI/ML engineer with production computer vision and generative AI experience at Mercedes-Benz R&D. Outside of that, I build small projects and actually benchmark them instead of just demoing them.
Gesture recognition for autonomous parking, knowledge distillation cutting inference time 8.5x, camera calibration and a 4-camera bird's-eye-view pipeline.
Worked on ExtractIQ, a document AI platform combining VLM and NLP pipelines across 200+ documents; trained YOLOv11 for region detection, deployed on Azure.
Deployed YOLO on Jetson AGX Orin with TensorRT quantization, cutting inference latency from ~350ms to sub-100ms.
I built a tool-calling harness for LLM agents around one question: when a tool breaks, does it recover, fail honestly, or fail silently? I proved it against six real failure modes instead of just demoing the happy path.
I built a drop-in wrapper that intercepts every LLM call and runs it through a measured token-reduction pipeline: caching, adaptive limits, compression, all without degrading the response.
I streamed real-time BPM from a pulse sensor while playing two games with very different pacing, then forecast it with a pretrained attention-based transformer.
I'm an early-career engineer working across computer vision, generative AI, and predictive modeling, from production systems in a car at Mercedes-Benz to small side projects I actually measure instead of just demo.
If something doesn't work yet, I'd rather say that plainly than dress it up. Everything else here is proven, not claimed.