
Researcher
Senior Data Scientist at Target
I study how reinforcement learning agents generalize to unseen environments — and whether the evaluation protocols the field relies on actually measure what they claim to.
I graduated from IIT Jammu with a B.Tech in Mechanical Engineering and a CS minor (Director's Gold Medal, 9.1/10 CGPA). My first research experience — on multi-scale feature extraction for image quality assessment — introduced me to the question of how architectural design choices affect model robustness, a thread I now pursue in the reinforcement-learning generalization setting.
After working as a data scientist — formulating real-world allocation and bidding problems as sequential decision-making under constraints — I deepened my RL foundations through independent study — working through Sutton & Barto alongside the NPTEL reinforcement learning lectures and Stanford's CS234 — and through the ISI Kolkata Winter Schools on Deep Learning (2023 and 2025). I am now pursuing independent research on evaluation methodology for RL generalization in procedurally generated environments, targeting Fall 2027 PhD admissions.
My research focuses on generalization in reinforcement learning — specifically in procedurally generated environments, where agents must transfer learned behavior from training levels to unseen test levels.
I am particularly interested in the evaluation methodology that underlies generalization claims in RL. How do we know a reported generalization gap reflects genuine transfer failure rather than an artifact of how we measured it? Standard practices often go unverified, which undermines our ability to make reliable scientific claims about architectural or algorithmic improvements.
Three threads
How choices in train-vs-test evaluation procedures can inflate or deflate measured generalization gaps, and how to build rigorous, matched-protocol evaluation the field can adopt.
A generalization gap is only meaningful if the policy being evaluated has actually learned — verifying policy convergence before interpreting any train–test comparison.
How encoder design and representation quality interact with policy learning to shape generalization.
Previously, I worked on multi-scale visual representations for image quality assessment, developing a Siamese architecture combining parallel feature-extraction blocks with transformers — competitive in the NTIRE 2022 IQA Challenge (CVPR Workshop, 7th place).
Pursuing independent research on evaluation methodology for RL generalization in procedurally generated environments — manuscripts under review.
Joined Target as a Senior Data Scientist, working on large-scale fulfilment-network optimization.
Selected for the Winter School on Deep Learning at ISI Kolkata.
Began research on transformer-based intrusion detection as a Research Engineer at ISRDC, IIT Bombay.
Preprint “Multi-Scale Features and Parallel Transformers Based Image Quality Assessment” released. arXiv ↗
Co-authored the NTIRE 2022 Challenge report on Perceptual Image Quality Assessment (CVPR Workshop). Paper ↗
Investigating evaluation methodology and reproducibility in reinforcement learning benchmarks, with a focus on identifying and correcting measurement artifacts that can inflate or deflate reported generalization gaps.
Exploring whether contrastive and self-predictive auxiliary losses (e.g., CURL, SPR, reconstruction objectives) improve policy generalization in procedurally generated environments when combined with standard PPO training. This project asks whether representation-learning objectives known to improve sample efficiency also improve zero-shot generalization to unseen levels when evaluated under a controlled protocol — connecting the auxiliary-loss literature to the evaluation methodology questions motivating the rest of this research program.
Developed a Siamese network combining parallel CNN and transformer feature-extraction blocks for no-reference image quality assessment. Achieved 7th place in the NTIRE 2022 Perceptual IQA Challenge at CVPR. Studied the effect of scaling factors, backbone architectures, loss functions, and data-augmentation strategies on IQA performance.
Peer-reviewed papers and preprints. Author name in bold.
One manuscript under review at TMLR, two under review at workshops, and one additional manuscript under review. Preprints forthcoming.
Abhisek Keshari, Komal, Sadbhawna, Badri Subudhi, et al.
This paper reports on the NTIRE 2022 challenge on perceptual image quality assessment (IQA), held in conjunction with the New Trends in Image Restoration and Enhancement workshop (NTIRE) at CVPR 2022. The challenge addre…
Abhisek Keshari, Komal, Sadbhawna, Badri Subudhi
This paper focuses on image quality assessment in multimedia content, using transformers and a multi-scale feature-extraction approach to account for distortions at different scales. A Siamese network is established betw…
Evaluation methodology for RL generalization in procedurally generated environments. Code audit of twelve public ProcGen codebases; corrected evaluation protocol combining deterministic action selection with entropy-based convergence diagnostics.
IISc Bangalore
Model approximation in MDPs and POMDPs. Taught by Prof. Aditya Mahajan (McGill / Mila).
Target Corporation
Formulating large-scale supply-chain optimization (item-level demand allocation across fulfilment centres) as linear programs with operational and capacity constraints.
ISI Kolkata
Advanced coursework in gradient-based optimization, deep-learning theory, semi-supervised and few-shot learning, explainable AI, and adversarial learning.
Target Corporation
Formulated large-scale display-advertising budget allocation as a reinforcement learning problem — sequential decision-making with policy optimization under budget and delivery constraints across 11 product categories. Integrated deep-learning response prediction with constrained optimization (LP/MILP solvers), and built a parallel training and data-preparation pipeline for RL experimentation, achieving a 4× reduction in training time.
ISRDC, IIT Bombay
Conducted a literature survey on transformer architectures for intrusion-detection systems. Designed cross-domain experiments applying transformer attention to graph-structured network data. Implemented LeViT-based architectures for malware classification, reducing misclassification and inference time on the MaleVis dataset.
ISI Kolkata
Coursework in deep-learning foundations, optimization, and neural network architectures — the first of two ISI Kolkata winter programs.
MeraPashu360
Formulated workload distribution as a Mixed-Integer Programming problem using Google OR-Tools. Built a warehouse-location optimization framework integrating demand estimation and geospatial modeling (hexagonal tessellation) with MIP-based optimization.
Mechanical Engineering · CS Minor
Mechanical Engineering (Major), Computer Science (Minor). CGPA 9.1/10. Director's Gold Medal for best all-round performance.
ISRDC, IIT Bombay · 2023
A comprehensive overview of transformer architectures and their applications in computer security, highlighting recent advancements and future research directions.
ISRDC, IIT Bombay · 2023
Exploring the integration of multi-scale feature extraction with parallel transformer models to enhance the accuracy and robustness of image quality assessment techniques.
Indian Institute of Technology, Jammu · Oct 2022
Director's Gold Medal is given to the student for having the best all-round performance amongst all the graduating students.
Indian Institute of Technology, Jammu · Oct 2022
Institute Silver Medal is given to the student for having the best academic performance (the highest CGPA) amongst the graduating students in the department.
Indian Institute of Technology, Guwahati · Sep 2021
Awarded Silver Medal at Inter IIT Tech Meet 2021 for outstanding performance in technical competitions.