avatar

Kushal Kumar

Senior Applied Scientist
Amazon
kushlku[at]amazon[dot]com


About Me

I am a Senior Applied Scientist at Amazon with 7+ years of applied AI/ML research experience for solving real-world problems. I hold a Bachelor’s degree in Mathematics and Scientific Computing from Indian Institute of Technology, Kanpur (IITK) batch of 2018, with a minor degree in Machine Learning and Applications. I am a recipient of General Proficiency Award for graduating with the highest CGPA of 8.9/10.0 (3.68/4.0) in my department alongside Special Appreciation Award, from Games and Sports council for outstanding contributions to sporting community in college.

My current research interests include:

Publications

  1. Anonymous (fourth author)


  2. CIKM 2023
    Kushal Kumar, Tarik Arici, Tal Neiman, Jinyu Yang, Shioulin Sam, Yi Xu, Hakan Ferhatosmanoglu, and Ismail Tutar
    CIKM '23: Proceedings of the 32nd ACM International Conference on Information and Knowledge Management

  3. CIKM 2022
    Kushal Kumar and Anoop Saladi
    CIKM '22: Proceedings of the 31st ACM International Conference on Information & Knowledge Management

  4. KDD 2021 TrueFact Workshop
    Tarik Arici, Kushal Kumar, Hayreddin Çeker, Anoop S V K K Saladi and Ismail Tutar
    Third International Workshop on Truth Discovery and Fact Checking: Making a Credible Web for Tomorrow

Salient Projects

Demeter: End-to-End GenAI Entity Resolution Workflow, 2024-present

Developed a scalable pay-as-you-go entity resolution workflow for products that integrates representation learning, entity matching, and generative AI to improve de-duplication pipeline. Key innovations include: (1) applying nonparametric clustering as a lossless alternative to blocking to scale pairwise matching by cluster merging; (2) developing a performant, multimodal distilled LLM matcher to estimate match probabilities; (3) leveraging efficient off-the-shelf LLM inference for cluster canonicalization; and (4) pruning the cluster–cluster graph to obtain the most informative edges for human feedback. The new workflow was provably more efficient and achieved 99% precision and at least 90% recall with more than reduction in human touchpoint.

Mergen: Scalable Product Retrieval and Blocking Service, 2023-2024

Built a scalable product retrieval service using multimodal CLIP-like models trained in a distributed setup. Optimized for large-batch training and efficient image loading, the models exhibited strong generalization across diverse product categories and product relationships, enabling task-agnostic retrieval and blocking for large-scale pairwise workflows. The service significantly outperformed baselines such as token-overlap based blocking system and achieved 95% recall at 18% better reduction ratio with multimodal retrieval capabilities and 73% cheaper search. This work led to CIKM’24 publication titled Unsupervised Multimodal Representation Learning for High Quality Retrieval of Similar Products at E-commerce Scale as an oral paper.

PAVE: Improve Recall of Product Attribute Extraction Models using Reinforcement Learning, 2022-2023

Product attribute extraction models often suffer from low recall when product data has missing or noisy information. In this project, we developed a Reinforcement Learning based method to improve recall of any product attribute extraction model by leveraging information from neighbors of the given product. We formulated this problem as a first order Markov Decision Process and trained a Reinforcement Learning agent using Proximal Policy Optimization (PPO) using novel rewards to predict agent actions to choose the best attribute value from a ranked list of product neighbors. We show that our method outperforms several baselines to yield 10% higher recall at same precision without needing larger context window than the initial context window of the extraction model. This work was published in CIKM’22 titled PAVE: Lazy-MDP based Ensemble to Improve Recall of Product Attribute Extraction Models as an oral paper.

Multi-Attribute Extraction using Transformers, 2021-2022

Extended Convolutional Neural Network (CNN) based single-attribute entity extraction to multi-attribute entity extraction using a Transformer-based architecture that was more performant. This reduced the need to maintain separate models per product attribute and locale, creating an operationally efficient model-building pipeline and faster goal delivery. Moreover, Self-attention exhibited better context awareness for accurate multi-attribute extraction based on attribute relationships, for example knowledge of product type being a candy increases probability of ‘orange’ being a flavor attribute than color. Such relational priors proved critical when attributes share overlapping domains.

Modeling Price-Per-Unit for Consumables Products, 2020-2021

Price-per-unit information is a crucial factor in purchase decision, especially in consumables products like grocery and beauty products. In this project, we developed a scalable lightweight model to predict price per unit of products at scale using its textual attributes like title and description as well as categorical information from product taxonomy. We formulated fact extraction as a two-step question answering problem to first predict type of quantity, like weight, volume or count and then extract all the relevant quantity information using multi-span architecture. We show that our method outperforms the state-of-the-art approaches like BERT while being lightweight and scalable. This work was published in KDD 2021 Truefact Workshop titled Solving Price Per Unit Problem Around the World: Formulating Fact Extraction as Question Answering.

Timeline

  • JUL, 2025 : Promoted to Senior Applied Scientist
    Two impactful production solutions and three publications.
  • JUL, 2023 : Joined Amazon New York office
    Applied Scientist II in Amazon Selection and Catalog Systems (ASCS) team.
  • APR, 2020 : Joined Amazon India office
    Applied Scientist I in India Machine Learning (IML) team.
  • JUN, 2018 : Joined Goldman Sachs Services India Limited
    Analyst in Market Risk Division, Bengaluru office.
  • MAY, 2018 : Graduated from IIT Kanpur
    CGPA of 8.9/10 and a minor degree in Machine Learning and Application.
  • SEP, 2017 : Finished Research Internship in IBM Research Labs
    Implemented NASA's IFSDAF algorithm in Python under guidance of Jagabondhu Hazra.
  • JUL, 2016 : Finished CMI Research Internship in Abstract Algebra
    Studied commutative algebra concepts under supervision of Prof. Krishna Hanumanthu.
  • JUL, 2014 : Enrolled in Indian Institute of Technology, Kanpur
    Bachelor of Science four-year program in Mathematics and Scientific Computing.

Hobbies

Beyond my academic and professional endeavors, I take delight in life’s simpler joys.

🏏 Cricket


📚 Reading & Intellectual Exploration



Powered by Jekyll and Minimal Light theme.