In my last post, I addressed why search terms used to cull data sets in discovery should not be protected as attorney work product. Today, I want to distinguish an attorney’s “investigative queries” (for case assessment, to hone searches or to identify privileged content) from “culling queries” (to generate data sets meeting a legal obligation, whether conceived by an attorney, client, vendor or expert). I contend culling queries warrant no work product protection from disclosure.
Let’s assume a producing party has a sizable collection of potentially responsive electronic information. Producing party concludes that it would be too costly, slow or unreliable to segregate the ESI by reading everything and, instead, decides to examine just those items that contain particular words or phrases. Keyword queries thus serve to divide the ESI into two piles: one that will be reviewed by counsel and another that no one and nothing will qualitatively review. The latter is the “discard pile.” Culling queries may be applied iteratively, first to collect data from the enterprise and later to cull the collection for review. The reductive process may entail the successive use of a client’s local and enterprise search capabilities and/or a law firm’s or vendor’s search tools.
The common thread is that each lexical search mechanism serves to exclude ESI lacking certain terms from substantive review. No one ever assesses the discards for relevance or responsiveness.
Now, if we could be confident that keyword culling worked reasonably well and that the persons who came up with keywords were lexical magicians, there’d be no need to worry over the discard pile. We could trust that what we don’t know doesn’t hurt us.
But we do know that a hefty slug of responsive items ends up in that discard pile. We know this because studies and experience have established that keyword search is a crude, mechanical filter. It leaves most of what we seek behind.
Whether we are leaving behind an endurable or unendurable volume of responsive items depends on just how poorly those keywords performed. To gauge that, we’ve got to know what queries were run. Continue reading






