Skip to content

StringMatchGrader: empty reference_response scores 1.0 for any response under prefix/suffix/substring/regex/contains_all/contains_any #199

Description

@shaurya416

Empty reference_response scores as a perfect match for prefix_match, suffix_match, substring_match, regex_match, contains_all and contains_any

openjudge/graders/text/_utils/string_match_compute.py:58-89, 93-124, 127-165, 169-206, 210-251, 255-293

def compute_prefix_match(reference: str, response: str, case_sensitive: bool = True, **kwargs):
    ...
    matched = cand.startswith(ref)          # True for every response when ref == ""
    ...

def compute_suffix_match(reference: str, response: str, case_sensitive: bool = True, **kwargs):
    ...
    matched = cand.endswith(ref)            # True for every response when ref == ""
    ...

def compute_regex_match(reference: str, response: str, pattern: str = "", case_sensitive: bool = True, **kwargs):
    pattern_str = pattern if pattern else reference   # "" -> empty regex, matches every response
    ...

def compute_substring_match(reference: str, response: str, case_sensitive: bool = False, bidirectional: bool = False, **kwargs):
    ...
    if bidirectional:
        matched = ref in cand or cand in ref  # also True when response == "" and ref is non-empty
    else:
        matched = ref in cand                 # True for every response when ref == ""
    ...

def compute_contains_all(reference: str, response: str, substrings=None, case_sensitive: bool = False, **kwargs):
    target_substrings = substrings if substrings else [reference]   # [""] -> always contained
    ...
# compute_contains_any has the same fallback

StringMatchGrader._aevaluate (openjudge/graders/text/string_match.py:143-148) defaults the gold answer to an empty string:

async def _aevaluate(self, reference_response: str = "", response: str = "", **kwargs):

It then passes that value straight to the compute function for the algorithm chosen at construction (string_match.py:180), without checking that it is non-empty. BaseGrader.aevaluate does not validate it either.

Measured

I copied the functions above verbatim into an isolated script (stdlib re/typing only, no repo import) and called them with each algorithm's default parameters from DEFAULT_PARAMS:

Case reference response matched score
Degenerate (empty ref): prefix_match, suffix_match, substring_match (bidirectional off and on), regex_match (pattern=""), contains_all / contains_any (substrings=None) "" "This is a completely unrelated, wrong, degenerate answer that should NOT match anything." True (all six algorithms) 1.0
Degenerate (empty response): substring_match, bidirectional=True "cat" "" True 1.0
Control A (must fail) "Hello" / "World" / "dog" "Goodbye world" / "Goodbye there" / "The cat sat on the mat" False 0.0
Control B (intended use) "Hello" / "World" / "cat" "Hello World" / "Hello World" / "The cat sat on the mat" True 1.0

With bidirectional=False, the same empty-response case correctly gives 0.0.

Consequence

Sometimes the gold answer arrives empty. That happens when a dataset field is present but blank, or when the key is missing from the row and the grader's default reference_response="" fills in (no mapper, or a callable mapper that defaults to ""). In that case StringMatchGrader reports score=1.0 with a "matched"/"found"/"Contains all 1 substrings" reason for any response, including a completely wrong one, under prefix_match, suffix_match, substring_match, regex_match, contains_all or contains_any.

With substring_match and bidirectional=True, an empty model response gets full credit against any reference.

Nothing in the result (no error, warning or flag in the metadata) tells these cases apart from a genuine match. Any aggregate score or RL reward built on this grader inherits the false positive.

A dict mapper whose path is missing yields None instead of "". The compute functions then raise TypeError rather than passing, so that path fails loudly.

The same file already guards against an empty reference for compute_word_overlap and compute_char_overlap (if len(ref_words) == 0: score = 0.0, if len(ref_chars) == 0: score = 0.0). The six functions above have no equivalent guard.

Suggested fix

  • Add the empty-reference guard that compute_word_overlap/compute_char_overlap already use to compute_prefix_match, compute_suffix_match and compute_substring_match: return 0.0 / matched=False, or return an error entry, which _aevaluate already maps to score 0.0, when reference is empty.
  • Apply the same guard to compute_regex_match when both pattern and reference are empty, and to compute_contains_all/compute_contains_any when they fall back to [reference] with an empty reference.
  • For bidirectional=True, do not count an empty response as contained in the reference.
  • Optionally, have StringMatchGrader._aevaluate reject or flag an empty reference_response for these algorithms unless pattern/substrings supplies the target.

The open analyzer-layer PR #198 proposes the same principle: a degenerate input should not silently produce a clean-looking verdict.

Happy to open the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions