[EDRM Editor’s Note: The opinions and positions are those of John Tredennick and Dr. William Webber. EDRM is grateful to Trusted Partner Merlin Search Technologies for permission to publish. Unless otherwise noted, all images are courtesy of Merlin Search Technologies.]
The most dangerous document in a case may be the one nobody knew how to search for.
It may describe a contract breach as “the supplier problem.” It may refer to a secret project by an undisclosed code name. It may contain a misspelled company name, an unexplained acronym or a seemingly harmless phrase whose significance becomes apparent only months later.
Keyword search will not find that document unless someone anticipates the words it contains.
That does not make keyword search obsolete. It makes keyword search incomplete. Legal teams can now combine it at scale with two methods that look for different signals: semantic search, which matches meaning, and continuous active learning, which learns relevance from reviewer decisions.

The question is no longer which method should replace the others. It is what becomes possible when all three work together.
Three methods, three blind spots
For decades, searching a large document collection meant supplying words and retrieving documents that contained them. The profession made that process increasingly sophisticated with stemming, proximity operators, fuzzy matching and Boolean logic. Modern lexical systems such as BM25 can also rank results, giving greater weight to distinctive terms and accounting for document length.
Keyword search earned its position. When an exact term is the signal, it is usually the best tool available. That includes:
- Statutory citations and quoted contract language.
- Case numbers, account numbers, docket identifiers and document IDs.
- Names, product codes and other distinctive terms.
- Court-ordered or negotiated search terms.
Keyword search is fast, transparent and easy to explain. A judge can understand a term list. That matters.
But keyword search has a structural limit: it finds the words the searcher thought of, spelled the way the searcher expected to find them.
Semantic search works differently. It compares the meaning of the query with the meaning expressed in the documents. A search for failure to perform might surface discussions of “material default,” “missed deliverables,” “unfulfilled obligations” or “the supplier problem,” even when those documents contain none of the words in the query.
Continuous active learning, or CAL, looks for a third kind of signal. As reviewers mark documents relevant or not relevant, a classifier learns from those decisions and ranks the remaining collection. It is not limited to the original query or to a general model of linguistic similarity. It learns what relevance looks like in this collection, for this matter, based on the team’s judgments.
Each method also has a characteristic weakness.
Keyword search can miss material expressed in unexpected language. Semantic search can overgeneralize and is generally weaker on exact identifiers. CAL depends on the quality and range of the examples reviewers provide; poor or unrepresentative judgments can send it in the wrong direction.
The failure modes differ enough to make the methods complementary. Methods that fail in the same way are redundant. Methods that fail differently can catch one another’s mistakes.
Read full post
https://www.jdsupra.com/legalnews/keyword-search-isn-t-dead-searching-4356874/




