Research Portfolio

Tushar Anand

Incoming Bachelor's Thesis@MPI SWS | CS+Economics@BITS Pilani

I work on computer vision and multimodal deep learning, with an interest in LLMs and VLMs.

Publications

04 papers
2026 CVPR 2026 Findings
A2Z-10M+: Geometric Deep Learning with A-to-Z BRep Annotations for AI-Assisted CAD Modeling and Reverse Engineering
Pritham Kumar Jena, Bhavika Baburaj, Tushar Anand, Vedant Dutta, Vineeth Ulavala, Sk Aziz Ali
The largest compilation of 10 million multi-modal annotations and metadata for 1 million ABC CAD models, enabling unprecedented BRep learning with high-resolution meshes, 3D hand-drawn sketches, geometric and topological information, and textual captions.
2026 ICRA 2026 First Author
DenVisCoM: Dense Vision Correspondence Mamba for Efficient and Real-time Optical Flow and Stereo Estimation
Tushar Anand et al.
A novel Mamba block and hybrid architecture for joint, accurate, real-time estimation of optical flow and disparity. Addresses the three-way trade-off of inference speed, memory footprint, and accuracy in dense multi-view geometry tasks.
2026 ICRA 2026 First Author
DensePercept-NCSSD: Vision Mamba towards Real-time Dense Visual Perception with Non-Causal State Space Duality
Tushar Anand et al.
A non-causal Mamba block-based model for unified real-time optical flow and disparity estimation. Fuses pairwise input images within a non-causal selective state space, reducing inference times while maintaining high accuracy and low GPU memory footprint.
2025 ICASSP 2025 Equal Contribution
ViM-Disparity: Bridging the Gap of Speed, Accuracy and Memory for Disparity Map Generation
Maheswar Bora*, Tushar Anand*, Saurabh Atreya, Aritra Mukherjee, Abhijit Das
A Visual Mamba (ViM) based architecture that dissolves the speed–accuracy–memory trade-off for real-time disparity map generation, along with a novel joint performance measure for DMG evaluation.

Experience

May – Jul 2024
ML Intern
DeepTek AI
  • Trained ResNet and InceptionV3 models for lung disease detection from chest imaging.
  • Quantized model sizes using FP16/FP8 precision and applied QAT and post-training quantization on ResNet.
  • Evaluated models with sensitivity, specificity, and AUC-ROC metrics.
May 2023 – Jan 2024
Intern
WILP @ BITS
  • Built a Flask + Python platform for AI-powered exam creation — MCQ and subjective question generation with NLP/PyTorch.
  • Implemented automatic MCQ grading and AI-assisted subjective marking via semantic similarity scoring.

Projects

01 · Nov 2024 – Present
LLMs for Multiple Paradigm Causal Graphs
Conversation dataset tool for simulating therapeutic dialogues (patient–therapist). Local LLM pipelines with Flask backend, Firebase real-time storage, and context management.
LangChainHuggingFaceFirebaseFlask
02 · Jul – Nov 2024
Multimodal CAD Model Generation
Adapted generative CAD models for multi-modal input (LLMs, BERT, ViT, point cloud CNNs). Trained MAE for distillation and VAE for novel CAD generation from the DeepCAD latent space.
PyTorchViTVAEPoint Cloud
03 · GenAI Lead @ ACM BPHC
Machine Translation & Wall Defect Detection
Led LSTM / RNN / Transformer training for EN→FR/DE/ES translation. Also applied CNNs for wall defect detection from Kaggle dataset.
TransformersLSTMCNN
04 · Mars Rover Team @ BPHC
Autonomous Mapping & Object Detection
3D camera-based mapping and point cloud generation of surroundings. Curated dataset and trained YOLO-style detection model for tool recognition (hammer, spanner).
Point CloudsObject Detection3D Vision

Technologies

Languages
Python · C / C++ · SQL · Java · R · React
Deep Learning
PyTorch · HuggingFace · LangChain · Llama.cpp · Sklearn
Infrastructure
Git · Firebase · Streamlit · Flask · Conda · Numpy · Pandas
Research Areas
Dense Prediction · Mamba / SSMs · Multimodal ML · CAD / 3D Vision
Quantization
QAT · PTQ · FP16 · FP8 inference
Attended
3DVSS · IIIT Hyderabad (Jun 2025)

Contact

I'm open to research collaborations, internship opportunities, and general conversation about computer vision and ML systems.