Generative AI · Programming assessment

Research

Studying whether AI-assisted grading can be consistent, aligned with rubrics, and useful to the people making educational decisions.

CASCON 2025Peer reviewed

Published at CASCON 2025

Evaluating Generative AI for CS1 Code Grading: Direct vs Reverse Methods

Co-authored research comparing prompting strategies for AI-assisted programming assessment, with emphasis on grading consistency and alignment with human expectations.

Memon, A., Mohamed, A. · 2025

Research focus

Beyond “can the model grade?”

The more useful questions are where approaches disagree, how decisions align with rubrics, and when human review should take over.

01

Consistency

Compare grading behavior across methods and repeated model decisions.

02

Alignment

Measure whether outputs reflect rubric expectations and human judgment.

03

Intervention

Identify uncertainty and disagreement that should trigger human review.

Current direction

Building evaluation workflows people can reproduce.

The research continues through practical Python pipelines for model comparison, semantic analysis, thresholding, and structured review.

Next About