Klára Janoušková
PhD Student · Computer Vision & Machine Learning · CTU Prague
I am a fourth-year computer vision and machine learning PhD student at the Czech Technical University in Prague (CTU), supervised by professor Jiří Matas.
My current research focuses on spatial understanding for vision-language models, ImageNet-scale recognition benchmarks and reannotation, and reinforcement learning for video object segmentation. Previously, I worked on fine-grained species classification and biodiversity benchmarks, test-time adaptation for segmentation, AI-assisted labelling for civil infrastructure inspection, and scene text detection and recognition.
I teach labs for the Machine Learning and Pattern Recognition course at CTU and co-supervise several BSc/MSc students.
During my undergraduate studies I interned at the Technion (Chaim Baskin, Alex Bronstein), IBM Research Zurich (Mattia Rigotti, Ioana Giurgiu, Cristiano Malossi), and the CVC at UAB (Dimosthenis Karatzas, Lluis Gomez).
- Sep 2026 Our paper 'Doomed to Re-Annotate, Forever: The ImageNet Story' was accepted to NeurIPS 2026 as a spotlight!
- Aug 2026 We have released STRAP, a dataset of structured region–text annotations for 2M web images from a single MLLM pass. Check out the project page!
- Jun 2026 I was selected as an outstanding reviewer for CVPR 2026!
- Mar 2026 Our project on Efficient Vision-Language Models received funding from Toyota Motor Europe.
- Feb 2026 Our paper 'Multimodal Large Language Models as Image Classifiers' was accepted to CVPR 2026 Findings!
I am now interested in spatial understanding for VLMs — improving the spatial representation for VLMs and efficient VLMs, building on our prior work on context-aware object recognition and closed-form adaptation (Koo-Fu CLIP). We released STRAP, structured region–text annotations for 2M web images from a single MLLM pass.
The aim is complete validation set reannotation — an ongoing, soon to be published project. Building on our analysis of ImageNet flaws and VLM-based recognition methods.
Investigating reinforcement learning for learned memory control in the Segment Anything Model 2 (SAM2), with the goal of improving long-form video object segmentation by dynamically managing the memory bank.
All projects (including past work) on the Projects page · Publications on the Publications page.