An Automated Pseudocode Grading Framework with Explainability
Nipuni Jayathilake ⋅ Sandareka Wickramanayake
Abstract
Manual grading of student-written pseudocode is time-consuming, inconsistent, and hard to scale, as pseudocode lacks formal syntax and admits multiple logically equivalent solutions that resist test-case or syntax-based grading. Prompted LLMs also struggle with inconsistent rubric interpretation and unreliable assessment of open-ended answers. We present PseudoScore, a context-aware grading framework with a criterion-wise Grader and Explainer. Given an answer $A$, question $Q$, and rubric criteria $\{C_1, \ldots, C_n\}$, the Grader uses cross-attention over shared embeddings $E_Q, E_A, E_{C_i}$ to predict per-criterion scores, aggregated as $$\text{Grade} = \sum_{i=1}^{n} \text{MLP}(E_Q, E_A, E_{C_i})$$ while the Explainer generates evidence-grounded rationales verified by a Consistency Check Module. On our new PSA Dataset (3,324 graded student solutions across 32 questions), PseudoScore achieves an MAE of $1.4119$ and $R^2$ of $0.4924$, reducing MAE by $27.1\%$ over the strongest LLM baseline, with explanation quality reaching a BERTScore F1 of $0.79$.
Chat is not available.
Successful Page Load