Empty reference_response scores as a perfect match for prefix_match, suffix_match, substring_match, regex_match, contains_all and contains_any
openjudge/graders/text/_utils/string_match_compute.py:58-89, 93-124, 127-165, 169-206, 210-251, 255-293
def compute_prefix_match(reference: str, response: str, case_sensitive: bool = True, **kwargs):
...
matched = cand.startswith(ref) # True for every response when ref == ""
...
def compute_suffix_match(reference: str, response: str, case_sensitive: bool = True, **kwargs):
...
matched = cand.endswith(ref) # True for every response when ref == ""
...
def compute_regex_match(reference: str, response: str, pattern: str = "", case_sensitive: bool = True, **kwargs):
pattern_str = pattern if pattern else reference # "" -> empty regex, matches every response
...
def compute_substring_match(reference: str, response: str, case_sensitive: bool = False, bidirectional: bool = False, **kwargs):
...
if bidirectional:
matched = ref in cand or cand in ref # also True when response == "" and ref is non-empty
else:
matched = ref in cand # True for every response when ref == ""
...
def compute_contains_all(reference: str, response: str, substrings=None, case_sensitive: bool = False, **kwargs):
target_substrings = substrings if substrings else [reference] # [""] -> always contained
...
# compute_contains_any has the same fallback
StringMatchGrader._aevaluate (openjudge/graders/text/string_match.py:143-148) defaults the gold answer to an empty string:
async def _aevaluate(self, reference_response: str = "", response: str = "", **kwargs):
It then passes that value straight to the compute function for the algorithm chosen at construction (string_match.py:180), without checking that it is non-empty. BaseGrader.aevaluate does not validate it either.
Measured
I copied the functions above verbatim into an isolated script (stdlib re/typing only, no repo import) and called them with each algorithm's default parameters from DEFAULT_PARAMS:
| Case |
reference |
response |
matched |
score |
Degenerate (empty ref): prefix_match, suffix_match, substring_match (bidirectional off and on), regex_match (pattern=""), contains_all / contains_any (substrings=None) |
"" |
"This is a completely unrelated, wrong, degenerate answer that should NOT match anything." |
True (all six algorithms) |
1.0 |
Degenerate (empty response): substring_match, bidirectional=True |
"cat" |
"" |
True |
1.0 |
| Control A (must fail) |
"Hello" / "World" / "dog" |
"Goodbye world" / "Goodbye there" / "The cat sat on the mat" |
False |
0.0 |
| Control B (intended use) |
"Hello" / "World" / "cat" |
"Hello World" / "Hello World" / "The cat sat on the mat" |
True |
1.0 |
With bidirectional=False, the same empty-response case correctly gives 0.0.
Consequence
Sometimes the gold answer arrives empty. That happens when a dataset field is present but blank, or when the key is missing from the row and the grader's default reference_response="" fills in (no mapper, or a callable mapper that defaults to ""). In that case StringMatchGrader reports score=1.0 with a "matched"/"found"/"Contains all 1 substrings" reason for any response, including a completely wrong one, under prefix_match, suffix_match, substring_match, regex_match, contains_all or contains_any.
With substring_match and bidirectional=True, an empty model response gets full credit against any reference.
Nothing in the result (no error, warning or flag in the metadata) tells these cases apart from a genuine match. Any aggregate score or RL reward built on this grader inherits the false positive.
A dict mapper whose path is missing yields None instead of "". The compute functions then raise TypeError rather than passing, so that path fails loudly.
The same file already guards against an empty reference for compute_word_overlap and compute_char_overlap (if len(ref_words) == 0: score = 0.0, if len(ref_chars) == 0: score = 0.0). The six functions above have no equivalent guard.
Suggested fix
- Add the empty-reference guard that
compute_word_overlap/compute_char_overlap already use to compute_prefix_match, compute_suffix_match and compute_substring_match: return 0.0 / matched=False, or return an error entry, which _aevaluate already maps to score 0.0, when reference is empty.
- Apply the same guard to
compute_regex_match when both pattern and reference are empty, and to compute_contains_all/compute_contains_any when they fall back to [reference] with an empty reference.
- For
bidirectional=True, do not count an empty response as contained in the reference.
- Optionally, have
StringMatchGrader._aevaluate reject or flag an empty reference_response for these algorithms unless pattern/substrings supplies the target.
The open analyzer-layer PR #198 proposes the same principle: a degenerate input should not silently produce a clean-looking verdict.
Happy to open the PR.
Empty
reference_responsescores as a perfect match for prefix_match, suffix_match, substring_match, regex_match, contains_all and contains_anyopenjudge/graders/text/_utils/string_match_compute.py:58-89, 93-124, 127-165, 169-206, 210-251, 255-293StringMatchGrader._aevaluate(openjudge/graders/text/string_match.py:143-148) defaults the gold answer to an empty string:It then passes that value straight to the compute function for the algorithm chosen at construction (
string_match.py:180), without checking that it is non-empty.BaseGrader.aevaluatedoes not validate it either.Measured
I copied the functions above verbatim into an isolated script (stdlib
re/typingonly, no repo import) and called them with each algorithm's default parameters fromDEFAULT_PARAMS:pattern=""), contains_all / contains_any (substrings=None)""bidirectional=True"cat""""Hello"/"World"/"dog""Goodbye world"/"Goodbye there"/"The cat sat on the mat""Hello"/"World"/"cat""Hello World"/"Hello World"/"The cat sat on the mat"With
bidirectional=False, the same empty-response case correctly gives 0.0.Consequence
Sometimes the gold answer arrives empty. That happens when a dataset field is present but blank, or when the key is missing from the row and the grader's default
reference_response=""fills in (no mapper, or a callable mapper that defaults to""). In that caseStringMatchGraderreportsscore=1.0with a "matched"/"found"/"Contains all 1 substrings" reason for any response, including a completely wrong one, underprefix_match,suffix_match,substring_match,regex_match,contains_allorcontains_any.With
substring_matchandbidirectional=True, an empty model response gets full credit against any reference.Nothing in the result (no error, warning or flag in the metadata) tells these cases apart from a genuine match. Any aggregate score or RL reward built on this grader inherits the false positive.
A dict mapper whose path is missing yields
Noneinstead of"". The compute functions then raiseTypeErrorrather than passing, so that path fails loudly.The same file already guards against an empty reference for
compute_word_overlapandcompute_char_overlap(if len(ref_words) == 0: score = 0.0,if len(ref_chars) == 0: score = 0.0). The six functions above have no equivalent guard.Suggested fix
compute_word_overlap/compute_char_overlapalready use tocompute_prefix_match,compute_suffix_matchandcompute_substring_match: return0.0/matched=False, or return anerrorentry, which_aevaluatealready maps to score 0.0, whenreferenceis empty.compute_regex_matchwhen bothpatternandreferenceare empty, and tocompute_contains_all/compute_contains_anywhen they fall back to[reference]with an empty reference.bidirectional=True, do not count an empty response as contained in the reference.StringMatchGrader._aevaluatereject or flag an emptyreference_responsefor these algorithms unlesspattern/substringssupplies the target.The open analyzer-layer PR #198 proposes the same principle: a degenerate input should not silently produce a clean-looking verdict.
Happy to open the PR.