• Home
  • About
  • CRAIGBALL.COM
  • Disclaimer
  • Log In

Ball in your Court

~ Musings on e-discovery & forensics.

Ball in your Court

Tag Archives: E-Discovery

Drafting RFPs for Robots to Read

20 Monday Jul 2026

Posted by craigball in ai, E-Discovery, Law Practice & Procedure

≈ 2 Comments

Tags

ai, artificial-intelligence, chatgpt, discovery, E-Discovery, eDiscovery, ESI Protocols, generative-ai, LLM, Requests for Production, RFP, technology

Time flies: Two years ago, in a post here, I floated the proposition that if the other side is going to hand your requests for production to a large language model and let the machine decide what’s responsive, then let’s draft those requests with the machine in mind. Doug Austin was generous enough to amplify the idea a week later. Two years on, it’s no longer something to merely think about; producing parties are ceding first-pass relevance review to LLMs.  The request in front of the model is now a prompt whether we know it or not.

So, let’s make that prompt our own.

A large language model doesn’t share the experience a seasoned reviewer brings to “all documents touching or concerning the transaction.” Ambiguity that a human reviewer resolves by instinct becomes, for a model, a coin flip between over- and under-inclusion. If you want their AI to find what you need, then tell it how to discriminate.

Seen that way, an AI-aware request for production is doing one of two things. At a minimum, it lards the request with enough discrete elements—custodians, systems, defined terms, date ranges, document types, examples—that opposing counsel can hardly avoid building those elements into whatever prompt they feed their review platform. At best (insofar as the rules of procedure permit or human sloth promotes), it hands them a fully formed prompt: language so ready to roll that the path of least resistance is to paste it straight away. The first mode constrains by specificity; the second exploits the happy truth that a good prompt ready-made is a prompt somebody will be tempted to deploy. Either way, the benefits are the same: clarity over boilerplate, context over conclusion, defined terms over loose keywords, and illustrative examples that let the model pattern-match to the documents you need.

None of this is way out there. It’s just good drafting, made newly consequential because the first reader is now a machine.

Oh, Those Pesky Rules

The Federal Rules reward this technique. FRCP Rule 34(b)(1)(A) requires that a request “describe with reasonable particularity each item or category of items to be inspected.” Particularity and prompt-craft pull on the same oar: both esteem the concrete over the conclusory. An AI-aware request isn’t a departure from Rule 34; it’s Rule 34 taken seriously by a lawyer who recognizes a model is on the other end.

Cautions to Stay Inside the Guardrails

First, drafting a request is not the same as running the other side’s review.  The Sedona Principles, Third Edition, Principle 6, holds that “responding parties are best situated to evaluate the procedures, methodologies, and technologies appropriate for preserving and producing their own electronically stored information.” I’ve quarreled with Sedona Six before—competence is something a producing party should endeavor to earn, not something we should presume—but the principle still rears its ugly head to shield the tools the other side chooses. So, the line to walk is this: an AI-aware request steers relevance and particularity; it does not dictate the responding party’s platform. Write the request to tell the model, any model, what responsiveness looks like. Don’t write it so as to effectively tell opposing counsel which model to choose or how to configure it. The former is advocacy and fair game. The latter invites a well-founded objection.

Second, particularity is a gun that kicks as hard as it shoots. The more precisely we enumerate document types and search terms, the greater the risk that a producing party treats a careful list as the outer boundary of the request and withholds everything beyond it. Two years ago, I noted this language would be “unlikely to be embraced by counsel ever-apprehensive of framing a request too-narrowly,” and the worry remains. The fix is craftsmanship: pair concrete guidance with a stated purpose and, okay, keep your cherished “including but not limited to,” so your examples instruct the model without unduly shrinking the scope.

There’s a third point worth mentioning, because it proves prompts are becoming discovery objects in their own right. In Conservation Law Foundation, Inc. v. Shell Oil Co., No. 3:21-cv-00933 (D. Conn. May 18, 2026), a magistrate judge ordered production of the prompts an expert used to drive an AI tool, treating them as fair game for discovery into methodology. That case concerns an expert’s prompts, not the language of an RFP (and, frankly, I don’t think the judge got it right in the face of a stipulation between the parties); so don’t read too much into it.  Still, the decision signals that courts have started to regard AI prompts as part of the discovery record. The prompt-craft we bring to our requests and the prompts our adversaries feed their review platforms are drifting toward daylight.

How Does It Work?

Take a matter everyone remembers. In the Dominion Voting Systems’ defamation suit against Fox News—the case that settled for $787.5 million in April 2023—the fight turned on what people inside Fox knew about the falsity of the fraud claims their network kept airing. A conventional, human-oriented request in that case might read like this, and requests like it are served every day:

All Documents and Communications relating to Dominion Voting Systems, including any allegation of fraud, vote manipulation, algorithmic “vote switching,” or foreign influence involving Dominion’s products in connection with the November 2020 U.S. presidential election.

Does a senior associate knows what to do with that? An AI model told to sort a Fox custodian’s mailbox against it will drown—everything “relates to” Dominion in a case about Dominion! A fallback to keywords will be, at once, over- and under-inclusive.

Let’s rewrite the request as something a reviewer will be sorely tempted to paste straight into a review tool:

Identify and produce every Document and Communication—including emails, text and Signal messages, Slack messages, on-air scripts, booking notes, and drafts—in which any Fox News host, producer, booker, or executive discussed, doubted, questioned, promoted, or sought to substantiate the claim that Dominion Voting Systems’ machines or software switched, deleted, or altered votes, or were connected to Smartmatic, Venezuela, Hugo Chávez, or the “Kraken.” The purpose of this request is to surface each custodian’s internal knowledge of the truth or falsity of the on-air fraud allegations. Responsive material will typically originate with custodians including the hosts and executives identified in Schedule A, date from November 1, 2020 through the network’s 2021 on-air corrections, and use terms such as “Dominion,” “Powell,” “Giuliani,” “rigged,” “switch,” “Smartmatic,” “Chávez,” or “crazy.” An example of a responsive document is an internal text message in which a host privately derided the fraud claims as baseless while the network continued to air them.

Notice what the second version does and doesn’t do. It gives the model purpose, custodians, boundaries, defined terms, and an example—everything a prompt needs, and enough discrete hooks that the other side can’t build its own prompt without importing most of them. It says nothing about which review platform to buy or how to configure and train it.

Use the Tools

One last point: If we’re going to draft requests for a machine to read, we might as well let a machine do the heavy lifting. Today’s subscription-tier models—the paid versions of ChatGPT, Claude, Gemini, or ChatGPT, not the free tiers coasting on last year’s tech—handle this conversion with ease. Feed one your draft and the bare contours of the case and let it do the heavy lifting. The prompt need be no fancier than this:

You are an experienced litigator revising a request for production so an AI tool conducting first-pass relevance review will read it accurately. Rewrite the request below to add reasonable particularity: name the likely custodians and data sources, add a one-sentence statement of purpose, enumerate the pertinent document types, list the defined terms and search language, set a sensible date range, and give one example of a responsive document—without dictating the responding party’s review methodology or narrowing the request’s reach. Here’s the request: [paste].

As always, mind the confidentiality of whatever you paste, use an account that won’t train on your inputs, and read every word the machine spits back—but let it spare you the first draft.

The robots are reading our requests now. Shouldn’t we write for our real audience?

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

The EDRM Isn’t Broken; It’s Misunderstood.

18 Wednesday Mar 2026

Posted by craigball in Uncategorized

≈ 4 Comments

Tags

E-Discovery, eDiscovery, EDRM, forensics, investigations

Disclaimer: I serve as General Counsel of EDRM, but this message is mine alone: Nothing that follows speaks for EDRM or its leadership.

I recently received a marketing email that contained this gem: Organizations are “asking if EDRM is structurally prepared for investigations.”

Short answer: Yes. Obviously. Because the EDRM was never a structure to begin with.

That’s not a knock. That’s the point.

The Electronic Discovery Reference Model is–wait for it–a reference model—not a workflow, not a platform architecture, not an operational blueprint. A reference model is a conceptual framework that identifies the principal stages and relationships in a process. It doesn’t tell you how to do something; it maps what needs doing. Think of it as a compass, not a GPS turn-by-turn. It orients you. It doesn’t drive for you, and it ain’t broke.

The EDRM diagram—that familiar left-to-right ribbon of stages from Information Governance through Presentation—has never pretended otherwise. It emerged before we had a framework to talk sensibly about the conceptual components of exchanging ESI as evidence. It has always depicted a reference arc, not a rigid assembly line. Wise practitioners always understood that the stages overlap, iterate, and telescope depending on the matter. You don’t march from Identification to Collection to Processing like soldiers in formation. You loop, you backtrack, you run stages in parallel, recurse and iterate. The model accommodates all of that because it describes the territory, not the trail. You want to merge or collapse several stages in your preferred workflow? Go for it! The EDRM doesn’t proscribe that, just as cramming several stages into a single super-stage doesn’t do away with the need to complete the tasks the sub-stages describe.

So when someone asks whether EDRM is “structurally prepared for investigations,” the premise is the problem. They’re evaluating a map by asking whether it can carry luggage.

The shift toward “governed internal workflows where legal, security, and compliance operate from a shared investigation infrastructure” is a legitimate operational development. Organizations should be thinking about unified investigation infrastructure. But that’s a workflow conversation—a conversation about tooling, governance, access controls, and process design. It is emphatically not a conversation that requires or benefits from declaring the EDRM obsolete or unprepared.

The EDRM doesn’t compete with your investigation workflow. It informs it. The moment you need to think about what data sources to identify, how to preserve without spoliation risk, how to collect defensibly, how to process for review, how to analyze and produce—there’s the EDRM, as useful and orienting as it has ever been.

What the EDRM can’t do—and was never meant to do—is be your ticketing system, your case management platform, or your chain-of-custody log. If your investigation workflow is broken, that’s not the EDRM’s fault for failing to be software. It’s a planning failure for expecting a reference model to do a workflow’s job.

The schematic is fine (circa 2014). The thinking around it sometimes isn’t.

Use the EDRM for what it is: a durable, vendor-neutral conceptual foundation that helps you ask the right questions in the right sequence. Then build or buy whatever workflow infrastructure serves your organization’s needs. The two aren’t in tension unless you insist on making them so.

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

Still on Dial-Up: Why It’s Time to Retire the Enron Email Corpus

15 Friday Aug 2025

Posted by craigball in Computer Forensics, E-Discovery, General Technology Posts

≈ 11 Comments

Tags

corpora, E-Discovery, eDiscovery, Enron, ESI, forensics

Early this century, when I was gaining a reputation as a trial lawyer who understood e-discovery and digital forensics, I was hired to work as the lead computer forensic examiner for plaintiffs in a headline-making case involving a Houston-based company called Enron.  It was a heady experience.

Today, everywhere you turn in e-discovery, Enron is still with us. Not the company that went down in flames more than two decades ago, but the Enron Email Corpus, the industry’s default demo dataset.

Type in “Ken Lay” or “Andy Fastow,” hit search, and watch the results roll in. For vendors, it’s the easy choice: free, legal, and familiar. But for 2025, it’s also frozen in time—benchmarking the future of discovery against the technological equivalent of a rotary phone. Or, now that AOL has lately retired its dial-up service, benchmarking it against a 56K modem.

How Enron Became Everyone’s Test Data

When Enron collapsed in 2001 amid accounting fraud and market-manipulation scandals, the U.S. Federal Energy Regulatory Commission (FERC) launched a sweeping investigation into abuses during the Western U.S. energy crisis. As part of that probe, FERC collected huge volumes of internal Enron email.

In 2003, in an extraordinary act of transparency, FERC made a subset of those emails public as part of its docket. Some messages were removed at employees’ request; all attachments were stripped.

The dataset got a second life when Carnegie Mellon University’s School of Computer Science downloaded the FERC release, cleaned and structured it into individual mailboxes, and published it for research. That CMU version contains roughly half a million messages from about 150 Enron employees.

A few years later, the Electronic Discovery Reference Model (EDRM)—where I serve as General Counsel—stepped in to make the corpus more accessible to the legal tech world. EDRM curated, repackaged, and hosted improved versions, including PST-structured mailboxes and more comprehensive metadata. Even after CMU stopped hosting it, EDRM kept it available for years, ensuring that anyone building or testing e-discovery tools had a free, legal dataset to use. [Note: EDRM no longer hosts the Enron corpus, but for those who like hunting antiques, you may find it (or parts of it) at CMU, Enrondata.org, Kaggle.com and, no joke, The Library of Congress].

Because it’s there, lawful, and easy, Enron became—and regrettably remains—the de facto benchmark in our industry.

Why Enron Endures

Its virtues are obvious:

  • Free and lawful to use
  • Large enough to exercise search and analytics tools
  • Real corporate communications with all their messy quirks
  • Familiar to the point of being an industry standard

But those virtues are also the trap. The data is from 2001—before smartphones, Teams, Slack, Zoom, linked attachments, and nearly every other element that makes modern email review challenging.

In 2025, running Enron through a discovery platform is like driving a Formula One race car on cobblestone streets.

Continue reading →

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...
Follow Ball in your Court on WordPress.com

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 2,241 other subscribers

Recent Posts

  • Fifteen Years in Your Court August 21, 2026
  • The AI Protective Order Double Standard July 27, 2026
  • Drafting RFPs for Robots to Read July 20, 2026
  • A Refresh of the Annotated ESI Protocol May 1, 2026
  • Free at Last: Ditching TurboTax for FreeTaxUSA April 5, 2026

Archives

RSS Feed RSS - Posts

CRAIGBALL.COM

Categories

EDD Blogroll

  • eDiscovery Today (Doug Austin)
  • eDiscovery Journal (Greg Buckles)
  • E-Discovery Law Alert (Gibbons)
  • Minerva 26 (Kelly Twigger)
  • Corporate E-Discovery Blog (Zapproved )
  • Basics of E-Discovery (Exterro)
  • Illuminating eDiscovery (Lighthouse)
  • The Relativity Blog
  • GLTC (Tom O'Connor)
  • E-D Team (Ralph Losey)
  • Complex Discovery (Rob Robinson)
  • Sedona Conference
  • CS DISCO Blog

Admin

  • Create account
  • Log in
  • Entries feed
  • Comments feed
  • WordPress.com

Enter your email address to follow Ball in Your Court and receive notifications of new posts by email.

Website Powered by WordPress.com.

  • Subscribe Subscribed
    • Ball in your Court
    • Join 2,093 other subscribers
    • Already have a WordPress.com account? Log in now.
    • Ball in your Court
    • Subscribe Subscribed
    • Sign up
    • Log in
    • Report this content
    • View site in Reader
    • Manage subscriptions
    • Collapse this bar
Loading Comments...
%d