PhD researcher in Computer Science at the University of Wyoming, advised by Dr. Yaqoob Majeed.
I study reasoning in large language and vision-language models, with a focus on counting and relational reasoning about 3D objects in complex scenes, from single images to multi-view 3D data and video. My earlier work is on explainable AI and weakly supervised semantic segmentation.
- Spatial and relational reasoning in VLMs: counting under occlusion, metric distance, perspective-taking, and cross-frame integration in video
- 3D scene understanding: monocular depth, open-vocabulary perception, and 3D scene graphs for reasoning
- Reinforcement learning with verifiable rewards for multimodal reasoning (GRPO / R1-style training)
- Explainable AI and weakly supervised segmentation
| Project | Summary |
|---|---|
| spatial-scene-bench | Procedural 3D scenes with exact ground truth (depth, instance masks, camera poses, orbit videos) and 11 occlusion-aware question types for diagnosing counting and spatial reasoning in VLMs. Zero-shot Qwen2.5-VL-3B handles single-view relations (92%) but drops to 24% on cross-frame counting and fails at metric distance. |
| scene-graph-spatial-reasoner | Lifts open-vocabulary detection, SAM masks and metric depth into a 3D scene graph for neuro-symbolic counting and relational reasoning. Amodal object completion and floor-based depth-scale calibration make multi-view fusion robust to 10% depth-scale error. |
| spatial-r1-vlm | R1-style GRPO with verifiable rewards (exact counts, MRA for metric answers, strict reasoning format) to improve 3D reasoning in Qwen2.5-VL on images and video. |
- Instance-Guided Class Activation Mapping for Weakly Supervised Semantic Segmentation
A. Torabi, Y. Majeed, M. M. Rahman, S. Gaihre. arXiv:2509.12496, 2025. [paper] - Using Cartesian Genetic Programming Approach with New Crossover Technique to Design Convolutional Neural
Networks
A. Torabi, A. Sharifi, M. Teshnehlab. Neural Processing Letters (Springer), 2023. - Integrating Deep CNN Models for Multilingual Sign Language Recognition: A SignLink-Based Approach for Bengali
and English
N. H. Samin, M. I. Ferdous, R. A. Suity, et al. Research Square preprint, 2025.
- agricultural-dataset-combination: a 116k-image unified agricultural dataset for WSSS
- FeelAnyForce: reimplementation of ViT / Swin contact-force estimation from vision-based tactile sensors
- Model-Extraction-Attack-GNN: reimplementation of realistic model extraction attacks against graph neural networks
- XAIMethods: Grad-CAM, saliency, Integrated Gradients and LIME on image classifiers
Applied computer vision and vision-language model projects (ViT, CLIP, BLIP-2, YOLOv8, TrOCR, diffusion, video understanding) are listed under repositories.


