Visual intelligence is one of the main research directions at CILab. We investigate computational methods that enable machines to extract, represent and reason about information from images, video and multimodal data, with a particular focus on deep learning.
Our research spans image classification, object detection and segmentation, representation learning, graph-based models, generative AI and multimodal learning. We are interested not only in predictive performance, but also in how effective representations can be learned when data, annotations or computational resources are limited.
A growing direction concerns efficient and sustainable AI. We investigate resource-efficient architectures and learning strategies aimed at reducing computational and memory requirements during training and inference. This direction connects contemporary deep learning with a long-standing CILab interest in computational efficiency and is particularly relevant when intelligent models must operate on resource-constrained or edge devices.
We also investigate multimodal and knowledge-aware learning, combining visual information with text, structured knowledge and other data modalities. Graph-based representations, knowledge graphs and vision-language models provide complementary ways to enrich visual information with semantic and contextual knowledge.
Generative models represent another increasingly important direction, both as tools for synthesizing visual and multimodal content and as mechanisms for learning richer representations from complex data.
Across these topics, our goal is to develop visual intelligence methods that can move beyond controlled benchmarks and operate effectively under real-world constraints.