Generation, Not Selection: LLM Proposals for Certified Access-Control Policy Mining
Gautam Kumar ⋅ Ravi Sundaram ⋅ Shamik Sural
Abstract
Attribute-Based Access Control (ABAC) policy mining—recovering a compact, auditable policy from an access log—is posed as minimum set cover and solved greedily, which enjoys an $s^*(1+\ln m)$ guarantee. That guarantee is conditional on a precondition the literature never states: the candidate pool handed to greedy must contain the optimal rules. Generation, not selection, is the binding constraint, and existing generators are purely syntactic—they surface a rule only when enough log records witness it, so they fail on precisely the rare exception rules that fill a heavy-tailed policy. We replace the generator with a language model proposing rules semantically, from attribute names, and keep the combinatorial selector to carry the guarantee: every proposal is filtered against the observed denials, so soundness is independent of the proposal mechanism and an unreliable proposer can only waste candidates, never grant access. Over five seeds on each of two enterprise domains, adding semantic proposal to a syntactic miner dominates it on ten of ten runs—strictly better on seven—with near-optimal policies (9.2–9.6 rules against $s^\* = 9$, versus 11.2–14.6) and false grants cut four- to sevenfold. Replacing the syntactic miner rather than augmenting it wins only where witnesses are scarce. An anonymization ablation isolates the mechanism: opaque attribute names cost two thirds of the model's rule recovery while leaving the syntactic baselines provably unmoved, and cost nothing where witnesses are plentiful. Semantic knowledge and log evidence are substitutes.
Chat is not available.
Successful Page Load