Abhisek Keshari

Abhisek Keshari

Researcher

Senior Data Scientist at Target

I study how reinforcement learning agents generalize to unseen environments — and whether the evaluation protocols the field relies on actually measure what they claim to.

Background

About

I graduated from IIT Jammu with a B.Tech in Mechanical Engineering and a CS minor (Director's Gold Medal, 9.1/10 CGPA). My first research experience — on multi-scale feature extraction for image quality assessment — introduced me to the question of how architectural design choices affect model robustness, a thread I now pursue in the reinforcement-learning generalization setting.

After working as a data scientist — formulating real-world allocation and bidding problems as sequential decision-making under constraints — I deepened my RL foundations through independent study — working through Sutton & Barto alongside the NPTEL reinforcement learning lectures and Stanford's CS234 — and through the ISI Kolkata Winter Schools on Deep Learning (2023 and 2025). I am now pursuing independent research on evaluation methodology for RL generalization in procedurally generated environments, targeting Fall 2027 PhD admissions.

Focus

Research

My research focuses on generalization in reinforcement learning — specifically in procedurally generated environments, where agents must transfer learned behavior from training levels to unseen test levels.

I am particularly interested in the evaluation methodology that underlies generalization claims in RL. How do we know a reported generalization gap reflects genuine transfer failure rather than an artifact of how we measured it? Standard practices often go unverified, which undermines our ability to make reliable scientific claims about architectural or algorithmic improvements.

Three threads

Evaluation protocol design

How choices in train-vs-test evaluation procedures can inflate or deflate measured generalization gaps, and how to build rigorous, matched-protocol evaluation the field can adopt.

Convergence diagnostics

A generalization gap is only meaningful if the policy being evaluated has actually learned — verifying policy convergence before interpreting any train–test comparison.

Architecture & representation

How encoder design and representation quality interact with policy learning to shape generalization.

Previously, I worked on multi-scale visual representations for image quality assessment, developing a Siamese architecture combining parallel feature-extraction blocks with transformers — competitive in the NTIRE 2022 IQA Challenge (CVPR Workshop, 7th place).

Updates

News

  1. Pursuing independent research on evaluation methodology for RL generalization in procedurally generated environments — manuscripts under review.

  2. Joined Target as a Senior Data Scientist, working on large-scale fulfilment-network optimization.

  3. Selected for the Winter School on Deep Learning at ISI Kolkata.

  4. Began research on transformer-based intrusion detection as a Research Engineer at ISRDC, IIT Bombay.

  5. Preprint “Multi-Scale Features and Parallel Transformers Based Image Quality Assessment” released. arXiv

  6. Co-authored the NTIRE 2022 Challenge report on Perceptual Image Quality Assessment (CVPR Workshop). Paper

Research

Projects

Under review

Evaluation Methodology for RL Generalization

Investigating evaluation methodology and reproducibility in reinforcement learning benchmarks, with a focus on identifying and correcting measurement artifacts that can inflate or deflate reported generalization gaps.

PyTorchProcGenPPO (CleanRL)Weights & Biases
Under review · TMLR

Auxiliary Learning Objectives for RL Generalization

Exploring whether contrastive and self-predictive auxiliary losses (e.g., CURL, SPR, reconstruction objectives) improve policy generalization in procedurally generated environments when combined with standard PPO training. This project asks whether representation-learning objectives known to improve sample efficiency also improve zero-shot generalization to unseen levels when evaluated under a controlled protocol — connecting the auxiliary-loss literature to the evaluation methodology questions motivating the rest of this research program.

2022

Multi-Scale Feature Extraction for Image Quality Assessment

Developed a Siamese network combining parallel CNN and transformer feature-extraction blocks for no-reference image quality assessment. Achieved 7th place in the NTIRE 2022 Perceptual IQA Challenge at CVPR. Studied the effect of scaling factors, backbone architectures, loss functions, and data-augmentation strategies on IQA performance.

Selected work

Publications

Peer-reviewed papers and preprints. Author name in bold.

One manuscript under review at TMLR, two under review at workshops, and one additional manuscript under review. Preprints forthcoming.

2022
  1. CVPR 2022 Workshop (NTIRE)

    NTIRE 2022 Challenge on Perceptual Image Quality Assessment

    Abhisek Keshari, Komal, Sadbhawna, Badri Subudhi, et al.

    This paper reports on the NTIRE 2022 challenge on perceptual image quality assessment (IQA), held in conjunction with the New Trends in Image Restoration and Enhancement workshop (NTIRE) at CVPR 2022. The challenge addre…

  2. Preprint · arXiv

    Multi-Scale Features and Parallel Transformers Based Image Quality Assessment

    Abhisek Keshari, Komal, Sadbhawna, Badri Subudhi

    This paper focuses on image quality assessment in multimedia content, using transformers and a multi-scale feature-extraction approach to account for distortions at different scales. A Siamese network is established betw…

Timeline

Experience

Jul 2026 – Present

Independent Research

Evaluation methodology for RL generalization in procedurally generated environments. Code audit of twelve public ProcGen codebases; corrected evaluation protocol combining deterministic action selection with entropy-based convergence diagnostics.

Jul 2026

CNI Summer School

IISc Bangalore

Model approximation in MDPs and POMDPs. Taught by Prof. Aditya Mahajan (McGill / Mila).

Sep 2025 – Present

Senior Data Scientist

Target Corporation

Formulating large-scale supply-chain optimization (item-level demand allocation across fulfilment centres) as linear programs with operational and capacity constraints.

Jan 2025 – Mar 2025

Winter School on Deep Learning

ISI Kolkata

Advanced coursework in gradient-based optimization, deep-learning theory, semi-supervised and few-shot learning, explainable AI, and adversarial learning.

Mar 2024 – Sep 2025

Data Scientist

Target Corporation

Formulated large-scale display-advertising budget allocation as a reinforcement learning problem — sequential decision-making with policy optimization under budget and delivery constraints across 11 product categories. Integrated deep-learning response prediction with constrained optimization (LP/MILP solvers), and built a parallel training and data-preparation pipeline for RL experimentation, achieving a 4× reduction in training time.

Aug 2023 – Dec 2023

Research Engineer

ISRDC, IIT Bombay

Conducted a literature survey on transformer architectures for intrusion-detection systems. Designed cross-domain experiments applying transformer attention to graph-structured network data. Implemented LeViT-based architectures for malware classification, reducing misclassification and inference time on the MaleVis dataset.

Jan 2023 – Mar 2023

Winter School on Deep Learning

ISI Kolkata

Coursework in deep-learning foundations, optimization, and neural network architectures — the first of two ISI Kolkata winter programs.

Jul 2022 – Aug 2023

Data Scientist

MeraPashu360

Formulated workload distribution as a Mixed-Integer Programming problem using Google OR-Tools. Built a warehouse-location optimization framework integrating demand estimation and geospatial modeling (hexagonal tessellation) with MIP-based optimization.

Jul 2018 – Jun 2022

B.Tech, IIT Jammu

Mechanical Engineering · CS Minor

Mechanical Engineering (Major), Computer Science (Minor). CGPA 9.1/10. Director's Gold Medal for best all-round performance.

Recognition

Awards & Achievements

Director's Gold Medal

Indian Institute of Technology, Jammu · Oct 2022

Director's Gold Medal is given to the student for having the best all-round performance amongst all the graduating students.

Institute Silver Medal

Indian Institute of Technology, Jammu · Oct 2022

Institute Silver Medal is given to the student for having the best academic performance (the highest CGPA) amongst the graduating students in the department.

Silver Medal, Inter IIT Tech Meet 2021

Indian Institute of Technology, Guwahati · Sep 2021

Awarded Silver Medal at Inter IIT Tech Meet 2021 for outstanding performance in technical competitions.