Skip to content
View selfishout's full-sized avatar

Highlights

  • Pro

Block or report selfishout

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
selfishout/README.md

Ali Torabi

PhD researcher in Computer Science at the University of Wyoming, advised by Dr. Yaqoob Majeed.

I study reasoning in large language and vision-language models, with a focus on counting and relational reasoning about 3D objects in complex scenes, from single images to multi-view 3D data and video. My earlier work is on explainable AI and weakly supervised semantic segmentation.

Website · LinkedIn · Email

Research interests

  • Spatial and relational reasoning in VLMs: counting under occlusion, metric distance, perspective-taking, and cross-frame integration in video
  • 3D scene understanding: monocular depth, open-vocabulary perception, and 3D scene graphs for reasoning
  • Reinforcement learning with verifiable rewards for multimodal reasoning (GRPO / R1-style training)
  • Explainable AI and weakly supervised segmentation

Current work

Project Summary
spatial-scene-bench Procedural 3D scenes with exact ground truth (depth, instance masks, camera poses, orbit videos) and 11 occlusion-aware question types for diagnosing counting and spatial reasoning in VLMs. Zero-shot Qwen2.5-VL-3B handles single-view relations (92%) but drops to 24% on cross-frame counting and fails at metric distance.
scene-graph-spatial-reasoner Lifts open-vocabulary detection, SAM masks and metric depth into a 3D scene graph for neuro-symbolic counting and relational reasoning. Amodal object completion and floor-based depth-scale calibration make multi-view fusion robust to 10% depth-scale error.
spatial-r1-vlm R1-style GRPO with verifiable rewards (exact counts, MRA for metric answers, strict reasoning format) to improve 3D reasoning in Qwen2.5-VL on images and video.

Publications

  • Instance-Guided Class Activation Mapping for Weakly Supervised Semantic Segmentation
    A. Torabi, Y. Majeed, M. M. Rahman, S. Gaihre. arXiv:2509.12496, 2025. [paper]
  • Using Cartesian Genetic Programming Approach with New Crossover Technique to Design Convolutional Neural Networks
    A. Torabi, A. Sharifi, M. Teshnehlab. Neural Processing Letters (Springer), 2023.
  • Integrating Deep CNN Models for Multilingual Sign Language Recognition: A SignLink-Based Approach for Bengali and English
    N. H. Samin, M. I. Ferdous, R. A. Suity, et al. Research Square preprint, 2025.

Other research code

Applied computer vision and vision-language model projects (ViT, CLIP, BLIP-2, YOLOv8, TrOCR, diffusion, video understanding) are listed under repositories.

Pinned Loading

  1. agricultural-dataset-combination agricultural-dataset-combination Public

    A comprehensive project for combining multiple agricultural datasets into a unified format suitable for Weakly Supervised Semantic Segmentation (WSSS) applications

    Python 1

  2. XAIMethods XAIMethods Public

    Grad-CAM, saliency maps, Integrated Gradients and LIME applied to a pretrained image classifier

    Python 1

  3. FeelAnyForce FeelAnyForce Public

    Reimplementation of FeelAnyForce: ViT and Swin Transformer models estimating 3D contact force from vision-based tactile sensor images

    Python

  4. scene-graph-spatial-reasoner scene-graph-spatial-reasoner Public

    Metric 3D scene graphs from open-vocabulary detection, SAM and monocular depth for counting and spatial reasoning over images and video

    Python

  5. spatial-r1-vlm spatial-r1-vlm Public

    R1-style GRPO with verifiable rewards to improve 3D counting and spatial reasoning in Qwen2.5-VL (images and video)

    Python

  6. spatial-scene-bench spatial-scene-bench Public

    Procedural 3D scenes with exact ground truth (depth, masks, poses, orbit videos) for evaluating counting and relational reasoning in vision-language models

    Python