• Home
  • About
  • CRAIGBALL.COM
  • Disclaimer
  • Log In

Ball in your Court

~ Musings on e-discovery & forensics.

Ball in your Court

Tag Archives: ESI Protocols

The AI Protective Order Double Standard

27 Monday Jul 2026

Posted by craigball in Uncategorized

≈ 4 Comments

Tags

ai, artificial-intelligence, chatgpt, eDiscovery, ESI Protocols, generative-ai

I’ve spent forty-odd years watching the discovery process get weaponized, not always to protect legitimate interests, but often to price opponents out of the fight. I’ve seen it with discovery, privacy and security protocols, with TAR disputes and with forensic examination demands that cost more than the case is worth. Every new technology becomes a vector for cost-shifting dressed up as diligence. AI protective orders are the latest iteration, using the same playbook.

Consider two rulings handed down the same day this June in the Southern District of New York. Orechovesky v. BNY Administrative Services, LLC, No. 1:25-cv-08517, 2026 WL 1725149 (S.D.N.Y. June 15, 2026), requires a party receiving protected material to certify that any AI tool used to process it will maintain confidentiality, won’t expose materials to unauthorized third parties, won’t train on inputs, and will allow deletion at the conclusion of the case. Sensible requirements. I don’t quarrel with them. The Court also requires the LLM to operate “within a closed, private, limited, secure universe”—whatever that means—and lists Relativity’s aiR Platform, Westlaw’s CoCounsel, and Gemini for Google Workspace (Enterprise or Business editions) as acceptable, without saying whether those tools exemplify the standard or are simply agreed-upon exceptions to it.

The second case, Pujas v. BDO USA, P.C., 2026 WL 1724307 (S.D.N.Y. June 15, 2026), goes further: it bars uploading confidential material to any AI tool unless the platform is “enterprise-grade,” backed by a binding agreement prohibiting the provider from using the data for training or product improvement, a so-called DPA (for Data Processing Agreement).

Enterprise-grade?!? It’s not clear what that means nor is it settled in law; but the notion is gaining traction in CLE panels and proposed orders: the assumption that only expensive, purpose-built legal AI platforms can satisfy requirements that every major consumer AI platform already satisfies. That assumption is wrong on the technology, wrong on the contracts, and wrong on the policy.

Courts are already proving my point. In Morgan v. V2X, Inc., No. 25-cv-01991-SKC-MDB (D. Colo. Mar. 30, 2026), the District of Colorado adopted a similar standard and then candidly recognized, “that practically speaking, and in light of the current state of AI, this provision will (at least for now) bar the parties from using most, if not all, mainstream low-to-no cost AI to process Confidential Information.”  If other courts follow uncritically, it will do what every prior technology-gatekeeping effort has done: widen the gap between well-funded litigants and everyone else, while delivering no meaningful improvement in data security.  My hope is that this post will shed light on a distinction without a difference so as to not hinder the use of properly configured, ‘consumer grade’ AI for processing sensitive data.

Same Engine, Different Price Tag

The legal AI industry doesn’t advertise it, but all the legal AI products on the market run on the same small handful of foundation models including OpenAI’s GPT series, Anthropic’s Claude, Google’s Gemini and are hosted on the same cloud infrastructure, like Azure, AWS or Google Cloud. Whether you access that model through a $150,000-per-year legal AI platform or a $20-a-month ChatGPT Plus or Claude Pro subscription, your data hits the same servers, gets processed by the same chips and is governed by the same operational security posture at the infrastructure layer (the actual servers and networks where data is processed).

The legal AI vendors do add value through workflow design, legal-specific prompting, citation checking, integration with document review tools and such; but what the vendors don’t add is a fundamentally different security architecture at the AI. The model doesn’t know whether it’s being called by a BigLaw firm’s bespoke platform or by little ol’ me. The bytes don’t care about the price tag on the Application Programming Interface (API) wrapper.

Certainly, the default security settings aren’t identical across vendors. Anthropic’s consumer plans—Free, Pro, and Max—don’t train on your conversations unless you affirmatively opt in. OpenAI’s consumer ChatGPT does the opposite: it trains on your conversations by default, unless you affirmatively opt out under Data Controls in Settings.

If you’re relying on a consumer subscription to satisfy a protective order, choosing the correct setting isn’t optional housekeeping; it’s the ballgame. For tools accessing the raw API, training the model on client data is not a concern; both Open AI and Anthropic have defaulted to “no training” for years.

When the ‘enterprise grade’ vendor tells you that your data is protected by their SOC 2 Type II attestation (an independent audit confirming security controls are in place and working), contractual commitments, and zero-data-retention processing (meaning your inputs aren’t stored after the response is generated), they’re describing protections that flow from the foundation model provider’s policies. The ‘consumer grade’ user benefits from the same no-training setting and the same underlying infrastructure security — though not from a zero-data-retention agreement, which is a negotiated feature, not a default at any price tier. What a consumer account gets instead is a short, fixed retention window, which I’ll come back to. The vendor’s DPA doesn’t cause OpenAI or Anthropic to not train on inputs. It merely documents a practice those providers already follow for every API customer, because their Fortune 500 clients demand it and because training on customer data creates legal liability they don’t want.

The security isn’t in the markup. It’s in the infrastructure. And the infrastructure is shared.

What the Contracts Actually Say

Compare the enterprise DPAs and the consumer terms of service, and you’ll find the operative commitments converge: no sharing with third parties, deletion on request, and—for the training question, once you’ve checked the box covered above—no training on inputs either. “The DPA says it in forty pages, dressed up with indemnification clauses and liability caps. The terms of service say it in four paragraphs; but the words that matter are substantively identical.

The DPA adds audit rights, breach notification timelines, and deletion SLAs (service-level agreements specifying timelines for action). These sound important, but what do they mean in practice?

Audit rights. No law firm audits OpenAI. No law firm sends a forensic examiner in to audit Microsoft’s data centers. The “audit right” exists to be pointed to in a certification, not to be exercised. It’s a clause that lets a general counsel tell a managing partner that the box is checked. Actual security verification comes from the provider’s SOC 2 report—which is public, and which I can read as easily as any Chief Information Security Officer.

Breach notification. The major providers must notify affected customers of a material breach regardless of contractual obligation, because many laws require it and the reputational cost of silence vastly exceeds the cost of disclosure. The contractual provision is belt-and-suspenders for something the provider’s self-interest already guarantees.

Deletion Service Level Agreements (SLA). Useful, genuinely. But consumer platforms also provide mechanisms to delete a chat or a project, even delete the account, if you open an account dedicated to a matter. The difference is that the enterprise customer gets a written confirmation within a defined timeframe. I get a confirmation screen. Both result in the same outcome at the infrastructure level.

I’m not saying the DPA is worthless. I’m saying it’s a contractual enhancement, not a technical one. And a protective order that requires a negotiated DPA as a precondition to AI use is requiring me to spend what I’ve seen run $10,000 to $30,000 for a solo or small-firm lawyer without existing DPA paper on file to produce a document that adds zero technical protection to discovery materials. That’s not proportionality. That’s a tax.

The Backup Hypocrisy

Let’s talk about “removal at conclusion,” because this is where the double standard gets embarrassing.

Go look at the return/destroy provision in standard protective orders and you won’t find much about backup media, disaster recovery replicas, or archived email. Courts didn’t carve out exceptions for such things because they never had to.

Here’s what happens at a large firm when a protective order requires return or destruction of discovery materials: someone decommissions the Relativity workspace. For old-school networks, the data persists on nightly backups for perhaps 30 to 90 days. It lingers in disaster recovery replicas (duplicate systems kept in case the primary goes down). It sits in email attachments that partners forwarded to associates who have since left the firm. It exists in litigation hold snapshots frozen for unrelated matters. Eventually—weeks or months later—the backup media cycles out, the replicas get overwritten and the data evaporates through the ordinary operation of retention schedules.

Courts didn’t require firms to forensically purge backup media simultaneously with the return/destroy deadline. The profession has long accepted this evaporating ‘digital tail,’ the unspoken understanding that transient, non-accessible, automatically expiring copies that no one will retrieve in the ordinary course will not be purged.  Why? Because those copies don’t create the risks that protective orders target. They can’t be readily searched, can’t be exploited competitively, and can’t be weaponized in other litigation. They’re ghosts in the machine, fading away in time.

A consumer AI platform’s 30-day compliance retention window is like that old backup media. The data exists, transiently, in a medium that isn’t accessible to me or anyone else, can’t be queried or searched, serves a narrow compliance function and expires automatically on a defined schedule. When I create a dedicated project for each matter—the AI equivalent of a separate Relativity workspace—and delete that project at the conclusion of the case when a protective order requires that I do so, I’ve done what BigLaw does when it decommissions a workspace. My 30-day tail is shorter and more predictable than their 90-day backup cycle.

If transient, non-accessible, automatically expiring retention didn’t violate a protective order when it lived on an Iron Mountain tape or, now, in a cloud backup, it doesn’t violate one when it lives on an Azure compliance server. Unless the standard is simply: the familiar gets a pass and the novel doesn’t.

The Solo Practitioner’s Structural Advantage

Here’s an irony that never gets acknowledged: my configuration is arguably more secure against unauthorized access than a typical firm deployment. Not less.

I don’t share my credentials. No one else has access to my account. No paralegal, no associate, no contract attorney logs in under my user ID or uses my workspace. When I process protected materials through a matter-specific project, the only human being who touches that data is me, bound by every ethical obligation the profession imposes. There’s a smaller attack surface to manage because there’s no team to manage.

A 200-lawyer firm deploying an enterprise AI platform across its litigation department must contend with dozens of credentialed users, role-based access controls (systems that limit what each user can see based on their assigned role), the ever-present risk that an associate working one case inadvertently queries materials from another, the chance that a lateral hire retains cached access after changing groups, and the unrelenting challenge of deprovisioning departing attorneys. Enterprise platforms address these problems with access control matrices, ethical walls and matter-segregation protocols, all of which add complexity. Complexity is where security fails.

I don’t have that problem. My isolation isn’t a workaround. It’s a happy accident of being solo, and it turns out to be superior to a multi-user deployment.

What Courts Should Reject

With that foundation laid, let me identify the provisions that courts should refuse to adopt—not because security doesn’t matter, but because these provisions don’t deliver real security, just excess cost. And cost without corresponding protection isn’t diligence. It’s a barrier to entry.

Mandated “enterprise-grade” platforms. If a court requires that AI tools be “enterprise-grade” or “specifically designed for legal use,” it’s requiring the same model on the same infrastructure, accessed through a more expensive wrapper. Courts should ask: what specific, technical security characteristic does the mandated tool possess that the alternative lacks? If the answer is “a negotiated DPA,” that’s a contractual characteristic, not a technical one.

Negotiated DPA requirements. A binding terms-of-service is an enforceable contract. It’s a contract of adhesion, sure, but so is every engagement letter a client signs with a litigation support vendor, and I’ve never seen a court require parties to negotiate bespoke terms with Microsoft or Relativity before loading discovery materials. If the operative terms address the substantive requirements of the protective order, the form of the contract shouldn’t matter.

Audit rights as a precondition. No one exercises audit rights over AI infrastructure. The solo practitioner and the BigLaw CISO rely on the same thing: the provider’s SOC 2 attestation and published security architecture. Requiring audit rights that will never be exercised is requiring a line item in a contract, not a security control.  It’s meaningless.

Blanket cloud-based prohibitions. If “processed on third-party infrastructure” is disqualifying for AI, it’s equally disqualifying for cloud-hosted document review, cloud-based legal research and every SaaS tool in the modern litigation stack. We decided years ago that cloud infrastructure with proper access controls provides adequate security for confidential materials. That conclusion doesn’t evaporate because the application performs inference rather than keyword search.

Absolute-guarantee certifications. No technology is breach-proof. No one certifies that Relativity will never be hacked, that a court reporter’s laptop is invulnerable, or that a lawyer won’t lose a laptop. We certify that reasonable precautions have been taken. The same standard should apply to AI tools. The formulation should track Rule 11: counsel certifies that, based on reasonable inquiry, the tool as configured satisfies each substantive requirement of the order.

Asymmetric restrictions. If AI is too dangerous for me to use with your client’s documents, it’s too dangerous for you to use with mine. Any AI restriction should apply reciprocally or not at all. Asymmetric technology limitations aren’t protective orders; they’re tactical weapons with a compliance veneer.

What Courts Should Require

I’m not arguing for anarchy or carelessness. I’m arguing for proportionality: the same principle that Federal Rule of Civil Procedure 26(b)(1) already applies to every other discovery burden, considering “the parties’ resources” and “whether the burden or expense of the proposed discovery outweighs its likely benefit.” If proportionality governs what a party must produce, it should equally govern what compliance infrastructure a party must procure to handle what it receives.

A properly scoped AI provision needs five things:

No training. The tool must not use inputs to improve its models. This is a binary setting available on every major platform at every subscription price tier (though as I noted above, on some platforms you switch it on, and on others you switch it off). Know which one you’re using, and check it.

No public accessibility. The tool must require authentication and must not expose one user’s inputs to another’s session. Every paid AI tool satisfies this by default.

Matter isolation. Discovery materials should be processed within a defined container—a project, workspace, or thread—that can be identified and deleted as a unit. This is how competent practitioners work regardless of what the protective order says.

Deletion at conclusion. The practitioner deletes the container and certifies deletion. Transient compliance retention gets the same grace we’ve always extended to backup media because it presents the same negligible risk profile.

Documentation. Counsel should be prepared to identify the tool, the configuration, and the contractual terms that address each requirement. Not a 40-page DPA. Just reasonable proof.

Five requirements, all achievable at any budget. All providing genuine protection against the actual risks that protective orders target: unauthorized use, competitive exploitation, and ongoing exposure. Anything beyond this isn’t really protecting data. It’s protecting market position.

The Stakes

I’ve been doing this long enough to know that our profession’s comfort with any technology follows a predictable curve: fear, restriction, grudging acceptance, ubiquity. We went through it with email, with cloud computing, with predictive coding. Each time, the early restrictions were driven by unfamiliarity rather than genuine risk. Each time, those restrictions disproportionately burdened smaller practitioners who couldn’t afford to buy their way past the gatekeepers.

AI is the most powerful leveling tool the legal profession has seen in my career. A solo practitioner with a well-configured AI tool can now perform work that previously required teams, like document analysis, deposition preparation and legal research at scale. That’s not a threat to justice; that’s the promise of justice. The small-firm lawyer handling a civil rights case on contingency, the public defender drowning in discovery, the solo practitioner taking on a corporate defendant with unlimited resources—these are the people AI helps most. And they’re the people who get locked out first when courts set compliance floors calibrated to BigLaw budgets, exactly as the Morgan court itself admitted its own standard might do.

We must not let that happen. Not because security doesn’t matter—it does—but because we can protect discovery materials without building a toll booth that only the well-heeled can pass through. In the ways that matter, the technology is the same. The commitments are the same. The security is the same. The only thing that differs is the price; and price has never been, and should never become, a proxy for diligence.

Hat tip to my friend Michael Berman, whose frequent and excellent series of posts about AI and discovery law got me thinking about this today.

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

Drafting RFPs for Robots to Read

20 Monday Jul 2026

Posted by craigball in ai, E-Discovery, Law Practice & Procedure

≈ 2 Comments

Tags

ai, artificial-intelligence, chatgpt, discovery, E-Discovery, eDiscovery, ESI Protocols, generative-ai, LLM, Requests for Production, RFP, technology

Time flies: Two years ago, in a post here, I floated the proposition that if the other side is going to hand your requests for production to a large language model and let the machine decide what’s responsive, then let’s draft those requests with the machine in mind. Doug Austin was generous enough to amplify the idea a week later. Two years on, it’s no longer something to merely think about; producing parties are ceding first-pass relevance review to LLMs.  The request in front of the model is now a prompt whether we know it or not.

So, let’s make that prompt our own.

A large language model doesn’t share the experience a seasoned reviewer brings to “all documents touching or concerning the transaction.” Ambiguity that a human reviewer resolves by instinct becomes, for a model, a coin flip between over- and under-inclusion. If you want their AI to find what you need, then tell it how to discriminate.

Seen that way, an AI-aware request for production is doing one of two things. At a minimum, it lards the request with enough discrete elements—custodians, systems, defined terms, date ranges, document types, examples—that opposing counsel can hardly avoid building those elements into whatever prompt they feed their review platform. At best (insofar as the rules of procedure permit or human sloth promotes), it hands them a fully formed prompt: language so ready to roll that the path of least resistance is to paste it straight away. The first mode constrains by specificity; the second exploits the happy truth that a good prompt ready-made is a prompt somebody will be tempted to deploy. Either way, the benefits are the same: clarity over boilerplate, context over conclusion, defined terms over loose keywords, and illustrative examples that let the model pattern-match to the documents you need.

None of this is way out there. It’s just good drafting, made newly consequential because the first reader is now a machine.

Oh, Those Pesky Rules

The Federal Rules reward this technique. FRCP Rule 34(b)(1)(A) requires that a request “describe with reasonable particularity each item or category of items to be inspected.” Particularity and prompt-craft pull on the same oar: both esteem the concrete over the conclusory. An AI-aware request isn’t a departure from Rule 34; it’s Rule 34 taken seriously by a lawyer who recognizes a model is on the other end.

Cautions to Stay Inside the Guardrails

First, drafting a request is not the same as running the other side’s review.  The Sedona Principles, Third Edition, Principle 6, holds that “responding parties are best situated to evaluate the procedures, methodologies, and technologies appropriate for preserving and producing their own electronically stored information.” I’ve quarreled with Sedona Six before—competence is something a producing party should endeavor to earn, not something we should presume—but the principle still rears its ugly head to shield the tools the other side chooses. So, the line to walk is this: an AI-aware request steers relevance and particularity; it does not dictate the responding party’s platform. Write the request to tell the model, any model, what responsiveness looks like. Don’t write it so as to effectively tell opposing counsel which model to choose or how to configure it. The former is advocacy and fair game. The latter invites a well-founded objection.

Second, particularity is a gun that kicks as hard as it shoots. The more precisely we enumerate document types and search terms, the greater the risk that a producing party treats a careful list as the outer boundary of the request and withholds everything beyond it. Two years ago, I noted this language would be “unlikely to be embraced by counsel ever-apprehensive of framing a request too-narrowly,” and the worry remains. The fix is craftsmanship: pair concrete guidance with a stated purpose and, okay, keep your cherished “including but not limited to,” so your examples instruct the model without unduly shrinking the scope.

There’s a third point worth mentioning, because it proves prompts are becoming discovery objects in their own right. In Conservation Law Foundation, Inc. v. Shell Oil Co., No. 3:21-cv-00933 (D. Conn. May 18, 2026), a magistrate judge ordered production of the prompts an expert used to drive an AI tool, treating them as fair game for discovery into methodology. That case concerns an expert’s prompts, not the language of an RFP (and, frankly, I don’t think the judge got it right in the face of a stipulation between the parties); so don’t read too much into it.  Still, the decision signals that courts have started to regard AI prompts as part of the discovery record. The prompt-craft we bring to our requests and the prompts our adversaries feed their review platforms are drifting toward daylight.

How Does It Work?

Take a matter everyone remembers. In the Dominion Voting Systems’ defamation suit against Fox News—the case that settled for $787.5 million in April 2023—the fight turned on what people inside Fox knew about the falsity of the fraud claims their network kept airing. A conventional, human-oriented request in that case might read like this, and requests like it are served every day:

All Documents and Communications relating to Dominion Voting Systems, including any allegation of fraud, vote manipulation, algorithmic “vote switching,” or foreign influence involving Dominion’s products in connection with the November 2020 U.S. presidential election.

Does a senior associate knows what to do with that? An AI model told to sort a Fox custodian’s mailbox against it will drown—everything “relates to” Dominion in a case about Dominion! A fallback to keywords will be, at once, over- and under-inclusive.

Let’s rewrite the request as something a reviewer will be sorely tempted to paste straight into a review tool:

Identify and produce every Document and Communication—including emails, text and Signal messages, Slack messages, on-air scripts, booking notes, and drafts—in which any Fox News host, producer, booker, or executive discussed, doubted, questioned, promoted, or sought to substantiate the claim that Dominion Voting Systems’ machines or software switched, deleted, or altered votes, or were connected to Smartmatic, Venezuela, Hugo Chávez, or the “Kraken.” The purpose of this request is to surface each custodian’s internal knowledge of the truth or falsity of the on-air fraud allegations. Responsive material will typically originate with custodians including the hosts and executives identified in Schedule A, date from November 1, 2020 through the network’s 2021 on-air corrections, and use terms such as “Dominion,” “Powell,” “Giuliani,” “rigged,” “switch,” “Smartmatic,” “Chávez,” or “crazy.” An example of a responsive document is an internal text message in which a host privately derided the fraud claims as baseless while the network continued to air them.

Notice what the second version does and doesn’t do. It gives the model purpose, custodians, boundaries, defined terms, and an example—everything a prompt needs, and enough discrete hooks that the other side can’t build its own prompt without importing most of them. It says nothing about which review platform to buy or how to configure and train it.

Use the Tools

One last point: If we’re going to draft requests for a machine to read, we might as well let a machine do the heavy lifting. Today’s subscription-tier models—the paid versions of ChatGPT, Claude, Gemini, or ChatGPT, not the free tiers coasting on last year’s tech—handle this conversion with ease. Feed one your draft and the bare contours of the case and let it do the heavy lifting. The prompt need be no fancier than this:

You are an experienced litigator revising a request for production so an AI tool conducting first-pass relevance review will read it accurately. Rewrite the request below to add reasonable particularity: name the likely custodians and data sources, add a one-sentence statement of purpose, enumerate the pertinent document types, list the defined terms and search language, set a sensible date range, and give one example of a responsive document—without dictating the responding party’s review methodology or narrowing the request’s reach. Here’s the request: [paste].

As always, mind the confidentiality of whatever you paste, use an account that won’t train on your inputs, and read every word the machine spits back—but let it spare you the first draft.

The robots are reading our requests now. Shouldn’t we write for our real audience?

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

A Refresh of the Annotated ESI Protocol

01 Friday May 2026

Posted by craigball in Computer Forensics, E-Discovery, Law Practice & Procedure

≈ 2 Comments

Tags

eDiscovery, ESI Protocols, generative-ai, Linked attachments

When I first published The Annotated ESI Protocol in January 2023, I hoped it would age the way most legal writing ages: slowly, with a few footnotes for the curious. Three years on, I’m releasing a substantially revised edition because the evidence has changed, and the protocols we negotiate need to change with it. The tools custodians use to communicate and collaborate in 2026 look materially different from just a few years ago, and the revised Annotated ESI Protocol seeks to address what we must accomplish today.

Download it here: http://www.craigball.com/Annotated_ESI_Protocol_2026_Final.pdf

The biggest changes are additions. Modern attachments—the cloud-hosted documents people send by pointers instead of as embedded files—now have their own section, with an exemplar provision distinguishing genuine substitute-attachments from documents merely referenced, a point-in-time-version obligation calibrated to what current collection tools can deliver, and a meet-and-confer trigger for cases where historical-version recovery genuinely matters. Short-message and collaboration-platform data—Slack, Teams, Google Chat—have a section requiring native export, a human-readable rendered transcript plus separately-produced attachments, with proportionality qualifiers acknowledging that not every producing party’s tooling produces all three out of the box. Mobile and ephemeral messaging get a tiered approach: consumer-grade backup-extraction utilities like iMazing are often acceptable for the run-of-the-mill civil cases, with forensic-grade collection (Cellebrite, AXIOM, and the rest) reserved for matters where alleged spoliation, deleted-content recovery, or device-integrity disputes warrant the greater rigor and cost. Audio, video, and voicemail get explicit native-production language. Search methodology, technology-assisted review, and generative AI in review each have their own short sections governing disclosure and validation. Foreign-language materials, which I had flagged in 2023 without addressing, now have an exemplar provision. In the years since I penned the original, we’ve benefited from insight on these topics gleaned from published case authority and thoughtful scholarship. The times are still a-changing and I’m reluctant to wade into contentious and unsettled topics; but at the same time, practitioners need guidance to move forward. Embrace what serves you and leave the rest.

A hybridized TIFF+ exemplar remains the primary form of production, because that’s what most institutional litigation still demands and a protocol that pretends otherwise won’t get traction, no matter the retrograde inefficiency of converting robust native formats to static images. But, the carve-outs for native production have grown to cover spreadsheets, presentations, databases, photographs, audio, video, short messages, mobile messages, structured-data exports, CAD, and anything else that doesn’t reduce sensibly to a static page image. If you can hold the line for native (instead of holding your nose for TIFFs), native is still better; the alternative language to get it is in there.

Addendum A addressing load file metadata production has expanded from twenty-seven fields to roughly sixty, organized into eleven labeled subsections covering the new evidence types, with explicit fields for collection-tool disclosure, modern-attachment metadata, conversation identifiers, edit and deletion flags, and platform tier. Thanks to the vigilance of my crack AI editor, a handful of drafting glitches in the 2023 edition are fixed.

The 2026 edition also acknowledges more directly than its predecessor that what counsel can demand and what a producing party’s platform and subscription tier can deliver are sometimes different things. The revised protocol asks more of producing parties—inevitably, considering all the new forms of ESI extant—but the new sections build in proportionality qualifiers, disclosure obligations and meet-and-confer triggers calibrated to what the leading platforms actually support in 2026. ESI protocols are still worth fighting for, and the better both sides understand their application and purpose, the less there is to bicker about.

If you’re interested in ESI Protocols and want to contribute your experience to an EDRM effort to frame a path to consensus protocols, please share your interest at https://edrm.net/edrm-projects/esi-protocol/

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

A Dog and Its Tail: Don’t Let Version Uncertainty Cloud Linked Attachment Production

02 Thursday Apr 2026

Posted by craigball in Computer Forensics, E-Discovery, Law Practice & Procedure

≈ 6 Comments

Tags

ESI Protocols, Linked attachments

Two years ago, I wrote a pair of posts (3/29/24 and 4/8/24) about linked attachments—what Microsoft calls “Cloud Attachments”—arguing that producing parties had been getting away with murder by not collecting and searching them.  The argument was straightforward: a linked attachment is no less relevant than an embedded one, the tools to collect them exist, and the claimed burdens were overstated.  Genuine, but exaggerated.

Nothing that’s happened since has changed that core proposition.  If anything, developments in case law, the Sedona Conference’s 2025 Commentary on collaboration platform discovery, and the emergence of proposed technical standards have reinforced it.  But those same developments carry a risk I want to flag: that the versioning question—which version of a linked attachment is the “right” one—is being elevated in ways that could hand producing parties a shiny new excuse for doing nothing.

What’s Changed in a Year

The landscape has shifted since, and largely in the right direction.

Courts are beginning to tiptoe towards what tools can actually do rather than accepting blanket claims of infeasibility.  The Carvana securities litigation is perhaps the most striking example: the court ordered a bounded forensic capability test using a specific tool, then expanded it when the initial pilot supported further testing.  That’s a different approach than we’ve seen before—a court saying, in effect, “show me what you can recover, don’t just tell me you can’t.”

The Sedona Conference published its Commentary on Discovery of Collaboration Platforms Data in 2025, acknowledging the distinct preservation, collection, and production challenges these platforms present.  When Sedona identifies a problem, that identification becomes part of the baseline against which “reasonable steps” under Rule 37(e) will be measured.  Parties who were aware of these challenges—and by now, every competent e-discovery practitioner should be—will find it increasingly hard to argue that their traditional, email-era workflow was good enough.

And a proposed technical standard—the Reconstruction-Grade eDiscovery Standard, authored by Peter Kozak and Brandon D’Agostino—has articulated an architectural framework for what preservation of collaborative evidence should look like.  It’s ambitious and thoughtful.  I want to engage with it constructively, because I think it gets several things right.  But I also want to sound a caution about how standards like this could be deployed in the real world of discovery disputes.

Two Problems

The RG standard does something valuable: it names and taxonomizes the specific ways that traditional preservation fails when evidence is collaborative, hyperlinked, and versioned.  Its framework identifies what it calls the “Preservation Gap” (the referenced content is never preserved at all) and the “Context Gap” (the content is preserved but not in the state it existed at the relevant time).  That’s a useful distinction.

But here’s where I part company—not with the standard’s laudable intent, but with the risk of how it may play out in the field.

The standard treats deterministic version resolution—preserving the as-sent version of a linked document, the version that existed when the message was transmitted—as a core conformance requirement.  Architecturally, I understand why.  If you’re building a system that aspires to reconstruction-grade fidelity, you want to capture the version the recipient would have seen when they clicked the link.  That’s the gold standard.

The problem is that the gold standard can become the enemy of any standard at all. 

To my eye, the versioning concern has been weaponized.  It goes like this: a requesting party asks for linked attachments.  The producing party raises the specter of versioning—“Which version do you want?  The as-sent version?  The as-accessed version?  The current version?  We can’t be sure which is the ‘right’ one, so the whole exercise is fraught with uncertainty.”  And that uncertainty becomes the justification for producing no version.  Not the wrong version.  No version.

That’s the tail wagging the dog.

The “Dog” Is Collection

The threshold obligation is to collect and search linked attachments.  Full stop.  A link in an email reveals nothing about the content of the linked document.  If you don’t collect the document, you can’t search it.  If you can’t search it, you can’t assess it for relevance.  And if you can’t assess it for relevance, you’re making a unilateral decision to exclude potentially responsive evidence—evidence that, but for a shift in how email systems handle large files, would have been embedded in the message and collected automatically.

That obligation exists independently of any versioning question.  It existed before anyone coined the term “reconstruction-grade.”  It existed when I wrote about it a year ago, and it existed for years before that.  “Perfect” is not the standard in e-discovery, but neither is “lousy.”

Beware, too, the half-measure.  A producing party, pressed on missing linked attachments, may offer to search the email text first and seek out the linked attachment only if the parent email hits on a keyword.  This sounds reasonable until you think about how email actually works.  It is exceedingly common for a transmitting email to say nothing more than “Please see attached” or “Here’s the draft we discussed,” while the attachment contains all the substantive content.  If the email text doesn’t trigger a keyword, the attachment—however rich in relevant material—never gets collected or searched.  And even if produced as a loose document, won’t tie to its “parent” transmitting message    

When we search email families containing embedded attachments, we treat the family as responsive if either the message or the attachment generates a hit.  Any workflow that conditions collection of linked attachments on hits in the transmitting email inverts that logic and guarantees that a large share of responsive evidence will be missed.

A producing party that collects and searches the current version of a linked attachment has done something meaningful.  They’ve brought the document into the review population.  They’ve assessed its content against the issues in the case.  They’ve preserved the family relationship between message and attachment.  They may not have captured the precise version that existed at send time, but they’ve captured a version—one that, in the overwhelming majority of cases, is likely to be the same or substantially similar to the transmitted version.

A producing party that collects nothing because of versioning uncertainty has done nothing.  Lousy.

The “Tail” Is Versioning

I don’t dismiss the versioning issue.  It’s real, and the RG standard is right to address it.  There are cases where the difference between the as-sent version and the current version matters enormously—a contract with terms that changed, a financial model with revised projections, a compliance policy that was softened after the relevant communication.  In those cases, producing the wrong version could mislead or, worse, could conceal what the actors actually relied upon.

But how often does this actually happen?

A year ago, I called for objective analysis: what percentage of cloud attachments are actually modified after transmittal?  I’m repeating the call, louder, because the industry still hasn’t answered it.

I have a strong intuition—and I want to be candid that it’s an intuition based on experience, not evidence—that the incidence of post-transmittal modification is modest overall.  My suspicion is that fewer than ten- to twenty percent of linked attachments are meaningfully modified after being shared, and perhaps far fewer than that.  Most cloud attachments are final or near-final documents shared for information, not living collaborative drafts.  Someone emails a report, a slide deck, a signed contract.  The link is a delivery mechanism, not an invitation to co-author.

But I also suspect the percentage varies widely depending upon the culture.  An organization whose culture runs to emailing finished work product will have a very different modification profile than one where teams routinely share early drafts via links for iterative editing in SharePoint.  A law firm circulating closing documents will look different from a product team sharing design specs that change daily.  The incidence of versioning concerns is likely a function of organizational work style, not some universal constant.

Here’s the point: I don’t have solid metrics.  I believe what I’m describing here, but belief is not evidence, and I would readily yield my suspicion to meaningful measurement.  The data needed to resolve this question is not exotic.  Any organization with a reasonably mature M365 environment could sample and compare the version history of linked attachments against the timestamps of the messages that transmitted them.  The analysis would tell us, for a given corpus, what percentage of linked attachments were modified after the transmitting message was sent, how significantly they were modified, and how soon after transmittal the modifications occurred.  That’s a study someone should do—a vendor, a consultant, an academic, a standards body.  It would replace speculation with evidence and give courts and practitioners a rational basis for calibrating the proportionality of versioning remediation.  Too, litigants coming to Court seeking relief from the duty to collect linked attachments should collect the metrics to measure the claimed risk and burden.

Until we have that data, we’re arguing about a problem whose magnitude we don’t grasp, while ignoring a problem whose magnitude is obvious: linked attachments aren’t being collected as they should be.

Don’t Throw Out the Baby

I want to be clear about what I’m not saying.  I’m not saying the RG standard is wrong to aspire to as-sent version resolution.  I’m not saying versioning doesn’t matter.  And I’m not attributing to the standard’s authors any intent to create a new excuse for non-production.  Reading the standard carefully, its concept of graduated conformance levels and its emphasis on proportionality suggest the opposite intent.

But standards exist in an adversarial ecosystem.  A standard that defines three conformance levels—RG-Core, RG-Plus, RG-Max—can be turned into a shield by a party arguing: “Your Honor, we can’t achieve even RG-Core conformance, so we shouldn’t be required to attempt collection of linked attachments.”  That argument confuses the standard’s aspirational architecture with the floor of a party’s discovery obligations.

The floor is not reconstruction-grade fidelity.  The floor is reasonable steps under Rule 37(e) and the obligation to search and produce relevant, responsive, non-privileged material.  That floor requires, at minimum, that you collect linked attachments using the tools your platform provides, search them, and produce responsive documents—even if you’re producing the current version rather than the as-sent version.

To put it another way: producing the “wrong” version of a responsive document is a problem.  Producing no version of a responsive document is a bigger problem.

I’ve been accused of leaning toward the interests of plaintiffs on this topic.  That’s neither fair nor accurate.  I advocate for evidence.  I’m committed to getting to the evidence that resolves disputes in what Rule 1 of the Federal Rules calls a “just, speedy, and inexpensive” fashion.  Not perfect.  Certainly not at any cost.  But I won’t accommodate high-handed, evasive approaches to the duty to produce responsive, non-privileged evidence—and dressing up a refusal to collect linked attachments in the language of versioning complexity is exactly that.

What the Standard Gets Right

Credit where it’s due.  Several elements of the RG framework strike me as genuinely constructive:

Exception transparency.  The standard requires structured records of what couldn’t be collected and why.  In the current landscape, failures are silent.  A linked attachment that can’t be retrieved simply disappears—no record that it was attempted, no record that it failed, no record of why.  Requiring a producing party to document its failures is a significant improvement over the status quo, where the absence of evidence is invisible.  Notably, courts have already begun requiring this kind of transparency on an ad hoc basis.  In the Uber litigation, Judge Cisneros ordered two custom metadata fields—“Missing Google Drive Attachments” and “Non-Contemporaneous”—to flag gaps and version discrepancies in the production.  What the RG standard proposes as a systemic architectural requirement, courts are already imposing case by case.  Formalizing that expectation is a natural and constructive next step.

The Preservation Gap vs. Context Gap distinction.  Naming these as separate failure modes is useful because they have different legal implications.  The Preservation Gap—evidence that was never preserved at all—maps cleanly to Rule 37(e).  The Context Gap—evidence preserved in the wrong state—is doctrinally murkier.  Courts don’t yet have a clean framework for “you preserved it, but what you preserved isn’t what was communicated.”  Distinguishing the two helps practitioners and courts think more precisely about what went wrong and what remedies are appropriate.

Capability testing as an emerging judicial norm.  The companion post to the standard highlights Carvana and the broader trajectory of courts ordering parties to demonstrate what their tools can do.  This is a welcome and overdue development.  The e-discovery conversation around linked attachments has too often been dominated by conclusory assertions of infeasibility.  Capability testing replaces assertion with demonstration, and that benefits everyone—including producing parties who have invested in the right tools and want credit for doing so.

Where We Go from Here

The path forward requires distinguishing between the immediate obligation and the aspirational architecture.

The immediate obligation  is collection.  If you’re on Microsoft 365, use Purview.  If you’re on Google Workspace, use Vault.  These tools aren’t perfect, but they exist, and they collect linked attachments.  The version you collect may be the current version rather than the as-sent version.  That’s a known limitation, not a reason to collect nothing.

The aspirational architecture  is reconstruction-grade fidelity—as-sent version resolution, deterministic exception handling, reproducible exports.  That’s where the industry needs to go.  Tools like Forensic Email Collector are already demonstrating that historical version recovery is technically possible in many cases.  The Carvana court’s willingness to order capability testing suggests that judges are ready to push the envelope.

But the bridge between those two isn’t “wait until perfect tools exist.”  The bridge is “do what you can now, document what you can’t, and improve your capabilities over time.”

That’s what proportionality actually means.  Not perfection.  Not paralysis.  But reasonable, good-faith efforts commensurate with the stakes and the state of the art.

The versioning problem will resolve because courts will order testing, because tools will improve, because someone will finally produce the empirical data on post-transmittal modification rates (pretty please), and because standards like the RG framework will mature.  These are all good-faith efforts to move the law and the industry forward, and they well deserve recognition for that commendable effort.

In the meantime, the producing party’s obligation is clear: collect the linked attachments, search them, and produce what’s responsive.

The tail does not get to wag the dog.

Hat tip to Doug Austin for highlighting the publication of the Reconstruction-Grade eDiscovery Standard on his eDiscovery Today blog.  Doug continues to be an indispensable resource for practitioners trying to keep pace with developments in this space.

© 2026 Craig D. Ball.  All rights reserved.

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

2026 Guide to AI and LLMs in Trial Practice

09 Friday Jan 2026

Posted by craigball in Uncategorized

≈ 2 Comments

Tags

ai, artificial-intelligence, chatgpt, eDiscovery, ESI Protocols, generative-ai, law, LLM

It’s been one year today since I published my introductory primer called Practical Uses for AI and LLMs in Trial Practice. AI changes so rapidly, I’ve been burning the midnight oil to overhaul and expand the work, now entitled Leery Lawyer’s Guide to AI and LLMs in Trial Practice. It’s no mere face lift, but a from-the-ground-up rewrite reflecting how AI and large language models power trial lawyer tasks today. Since the first edition, AI has moved from curiosity to necessity. Tools like ChatGPT and Harvey are no longer novelties, and the economics of AI-assisted drafting, discovery management, and record comprehension are undeniable. At the same time, the risks of use are better understood. Hallucinations, overreach, privilege exposure, and misplaced confidence are genuine, and the guide meets them head-on, offering practical guardrails and practice tips.

What’s new for 2026 is not more breathless talk of “transformation,” but a clearer picture of what works, what doesn’t, and what still demands adult supervision. The guide now speaks to lawyers who remain leery but are ready to use AI cautiously and competently. It expands beyond first forays to practical, defensible workflows: depositions, motion practice, ESI protocols, voir dire, and making sense of large records without losing the thread. It distinguishes consumer and enterprise tools, explains why governance matters, and emphasizes verification as a professional duty. Crucially, I cover the steps and prompts that get you going. If you’re looking for more hype, this isn’t it. If you want a practical field guide for using AI without surrendering judgment—or credibility—I hope you’ll take a look.

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

Garden Variety: Byte Fed. v. Lux Vending

12 Wednesday Jun 2024

Posted by craigball in E-Discovery

≈ 9 Comments

Tags

ESI Protocols, search

My esteemed colleagues, Kelly Twigger and Doug Austin, each posted about a recent discovery decision from the Middle District of Florida, case no. 8:23-cv-102-MSS-SPF, styled, Byte Fed., Inc. v. Lux Vending LLC. and decided by United States Magistrate Judge Sean Flynn on May 1, 2024.

Kelly and Doug share their customarily first-rate analyses of the ruling insofar as its finding that the assertion of boilerplate objections serves as a waiver.  The Court spanked defendant, The Cardamone Consulting Group, LLC, for its conduct.  That’s been picked apart elsewhere, and I have nothing to add.  I write here to address a feature of the dispute that no one has discussed (and sadly, neither did the Court), being the nature of the request for production that prompted the boilerplate objection of “vague and incomprehensible.”  We can learn much more from the case than just boilerplate=waiver.

Let’s look at the underlying request:

DOCUMENT REQUEST NO. 7:

All documents and electronically stored information that are generated in applying the search terms below to Your corporate email accounts (including but not limited to the email accounts for Nicholas Cardamone, Daniel Cardamone, and Patrick McCloskey):

ByteBitcoin w/s FloridaStanton
ByteFederalBitcoin w/s trademarkBranden w/3 Tawil
Byte FederallawsuitBrandon w/3 Mintz
most w/5 trustedScott w/3 BuchananDKI
Google w/s trademarkconfusion or confusedDynamic w/5 keyword

In its Motion to Compel, Plaintiff calls this request “clear on its face, and … a garden-variety type of request for production in connection with narrowly tailored search terms.”  The Plaintiff adds, “[y]et during the parties’ meet-and-confer, and although Cardamone’s counsel claimed that she was familiar with electronic discovery, the assertion was that her client – a company that has purportedly generated hundreds of millions of dollars in connection with online advertising and electronic data – ‘did not understand what to do.’”

So, Dear Reader, would you understand what to do? You’re steeped in electronic discovery—that’s why you’ve stopped by—but is the request clear, narrowly tailored and “garden-variety” such that we can apply it to a proper production workflow?  A few points to ponder:

1. There’s nothing in the Federal Rules of Civil Procedure that prohibits a request to run specific queries against databases, and email accounts are databases.  Rule 34 requires only that the request “describe with reasonable particularity each item or category of items to be inspected.” 

Conventional requests are couched in language geared to relevance; that is, the requests seek documents and ESI about a topic.  Counsel must then apply the law and the facts to guide clients in identifying responsive information.  Counsel reviews the information gathered and decides whether it’s responsive or should be withheld as a matter of right or privilege.

Over time, the notion took hold that sifting through electronically stored information was unduly burdensome, so opposing parties were expected to work together to fashion queries–“search terms” –to narrow the scope of review.  These keyword negotiations run the gamut from laughable to laudable. They’re duels between counsel frequently unarmed with knowledge of the search tools and processes or of the data under scrutiny.  In short, they use their ginormous lawyer brains to guess what might work if the digital world were as they imagine it to be.

Here, the plaintiff cuts to the chase, eschewing a request couched in relevance in favor of asking that specific searches be run: half of them Boolean constructs employing two types of proximity connectors. 

Was this smart?   You decide.

Continue reading →

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

Image

What’s All the Fuss About Linked Attachments?

29 Friday Mar 2024

Tags

ESI Protocols, hyperlinked files, Linked attachments, Purview

In the E-Discovery Bubble, we’re embroiled in a debate over “Linked Attachments.” Or should we say “Cloud Attachments,” or “Modern Attachments” or “Hyperlinked Files?” The name game aside, a linked or Cloud attachment is a file that, instead of being tucked into an email, gets uploaded to the cloud, leaving a trail in the form of a link shared in the transmitting message. It’s the digital equivalent of saying, “It’s in an Amazon locker; here’s the code” versus handing over a package directly.  An “embedded attachment” travels within the email, while a “linked attachment” sits in the cloud, awaiting retrieval using the link.

Some recoil at calling these digital parcels “attachments” at all. I stick with the term because it captures the essence of the sender’s intent to pass along a file, accessible only to those with the key to retrieve it, versus merely linking to a public webpage.  A file I seek to put in the hands of another via email is an “attachment,” even if it’s not an “embedment.” Oh, and Microsoft calls them “Cloud Attachments,” which is good enough for me.

Regardless of what we call them, they’re pivotal in discovery. If you’re on the requesting side, prepare for a revelation. And if you’re a producing party, the party’s over.

A Quick March Through History

Nascent email conveyed basic ASCII text but no attachments.  In the early 90s, the advent of Multipurpose Internet Mail Extensions (MIME) enabled files to hitch a ride on emails via ASCII encoded in Base64. This tech pivot meant attachments could join emails as encoded stowaways, to be unveiled upon receipt.

For two decades, this embedding magic meant capturing an email also netted its attachments. But come the early 2010s, the cloud era beckoned. Files too bulky for email began diverting to cloud storage with emails containing only links or “pointers” to these linked attachments. 

The Crux of the Matter

Linked attachments aren’t newcomers; they’ve been lurking for over a decade. Yet, there’s a growing “aha” moment among requesters as they realize the promised exchange of digital parcels hasn’t been as expected. Increasingly—and despite contrary representations by producing parties—relevant, responsive and non-privileged attachments to email aren’t being produced because relevant, responsive and non-privileged attachments aren’t being searched.

Wait! What?  Say that again.

You heard me.  As attachments shifted from being embedded to being linked, producing parties simply stopped collecting and searching those attachments.

How is that possible?  Why didn’t they disclose that? 

I’ll explain if you’ll indulge me in another history lesson.

Echoes From the Past

Traditionally, discovery leaned on indexing the content of email and attachments for quicker search, bypassing the need to sift through each individually.  Every service provider employs indexed search. 

When attachments are embedded in messages, those attachments are collected with the messages, then indexed and searched.  But when those attachments are linked instead of embedded, collecting them requires an added step of downloading the linked attachments with the transmitting message.  You must do this before you index and search because, if you fail to do so, the linked attachments aren’t searched or tied to the transmitting message in a so-called “family relationship.”

They aren’t searched.  Not because they are immaterial or irrelevant or in any absolute sense, inaccessible; a linked attachment is as amenable to being indexed and searched as any other document. They aren’t searched because they aren’t collected; and they aren’t collected because it’s easier to blow off linked attachments than collect them.

Linked attachments, squarely under the producer’s control, pose a quandary. A link in an email is a dead-end for anyone but the sender and recipients and reveals nothing of the file’s content. These linked attachments could be brimming with relevant keywords yet remain unexplored if not collected with their emails.

So, over the course of the last decade, how many times has an opponent revealed that, despite a commitment to search a custodian’s email, they were not going to collect and search linked documents?

The curse and blessing of long experience is having seen it all before.  Every generation imagines they invented sex, drugs and rock-n-roll, and every new information and communication technology is followed by what I call the “getting-away-with-murder” phase in civil discovery.  Litigants claim that whatever new tech has wrought is “too hard” to deal with in discovery, and they get away with murder by not having to produce the new stuff until long after we have the means and methods to do so.  I lived through that with e-mail, native production, then mobile devices, web content and now, linked attachments.

This isn’t just about technology but transparency and diligence in discovery. The reluctance to tackle linked attachments under claims of undue burden echoes past reluctances with emerging technologies. Yet, linked attachments, integral to relevance assessments, shouldn’t be sidelined.

What is the Burden, Really?

We see conclusory assertions of burden notwithstanding that the biggest platforms like Microsoft and Google offer ‘pretty good’ mechanisms to deal with linked attachments.  So, if a producing party claims burden, it behooves the Court and requesting parties to inquire into the source of the messaging.  When they do, judges may learn that the tools and techniques to collect linked attachments and preserve family relationships exist, but the producing party elected not to employ them.  Granted, these tools aren’t perfect; but they exist, and perfect is not the standard, just as pretending there are no solutions and doing nothing is not the standard. 

Claims that collecting linked attachments pose an undue burden because of increased volume are mostly nonsense.  The longstanding practice has been to collect a custodian’s messages and ALL embedded attachments, then index and search them.  With few exceptions, the number of items collected won’t differ materially whether the attachment is embedded or linked (although larger files tend to be linked).  So, any party arguing that collecting linked attachments will require the search of many more documents than before is fibbing or out of touch.  I try not to attribute to guile that which may be explained by ignorance, so let’s go with the latter.

Half Baked Solutions

Challenged for failing to search linked attachments, a responding party may protest that they searched the transmitting emails and even commit to collecting and searching linked attachments to emails containing search hits.  Sounds reasonable, right?  Yet, it’s not even close to reasonable. Here’s why:

When using lexical (e.g., keyword) search to identify potentially responsive e-mail “families,” the customary practice is to treat a message and its attachments as potentially responsive if either the content of the transmitting message or its attachment generates search “hits” for the keywords and queries run against them.  This is sensible because transmittals often say no more than, “see attached;” it’s the attachment that holds the hits.  Yet, stripped of its transmittal, you won’t know the timing or circulation of the attachment. So, we preserve and disclose email families.

But, if we rely upon the content of transmitting messages to prompt a search of linked attachments, we will miss the lion’s share of responsive evidence.  If we produce responsive documents without tying them to their transmittals, we can’t tell who got what and when.  All that “what did you know and when did you know it” matters.

Why Guess When You Can Measure?

Hopefully, you’re wondering how many hits suggesting relevance occur in transmittals and how many in attachments?  How many occur in both?  Great questions!  Happily, we can measure these things.  We can determine, on average, the percentage of messages that produce hits versus their attachments. 

If you determine that, say, half of hits were within embedded attachments, then you can fairly attribute that character to linked attachments not being searched.  In that sense, you can estimate how much you’re missing and ascertain a key component of a proper proportionality analysis.

So why don’t producing parties asserting burden supply this crucial metric? 

The Path Forward

Producing parties have been getting away with murder on linked attachments for so long that they’ve come to view it as an entitlement. Linked attachments are squarely within the ambit of what must be assessed for relevance.  The potential for a linked attachment to be responsive is no less than that of an item transmitted as an embedded attachment.  So, let’s stop pretending they have a different character in terms of relevance and devote our energies to fixing the process.

Collecting linked attachments isn’t as Herculean as some claim, especially with tools from giants like Microsoft and Google easing the process. The challenge, then, isn’t in the tools but in the willingness to employ them.

Do linked attachments pose problems?  They absolutely do!  I’ve elided over ancillary issues of versioning and credentials because those concerns reside in the realm between good and perfect solutions. Collection methods must be adapted to them—with clumsy workarounds at first and seamless solutions soon enough.  But in acknowledging that there are challenges, we must also acknowledge that these linked attachments have been around for years, and they are evidence.  Waiting until the crisis stage to begin thinking about how to deal with them was a choice, and a poor one.  I shudder to think of the responsive information ignored every single day because this issue is inadequately appreciated by counsel and courts.

Happily, this is simply a technical challenge and one starting to resolve.  Speeding the race to resolution requires that courts stop giving a free pass to the practice of ignoring linked attachments.  Abraham Lincoln defined a hypocrite as a “man who murdered his parents, and then pleaded for mercy on the grounds that he was an orphan.”  Having created the problem and ignored it for years, it seems disingenuous to indulge requesting parties’ pleas for mercy.  

In Conclusion

We’re at a crossroads, with technical solutions within reach and the legal imperative clearer than ever. It’s high time we bridge the gap between digital advancements and discovery obligations, ensuring that no piece of evidence, linked or embedded, escapes scrutiny.

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...

Posted by craigball | Filed under Computer Forensics, E-Discovery, Uncategorized

≈ 18 Comments

The Annotated ESI Protocol

09 Monday Jan 2023

Posted by craigball in Computer Forensics, E-Discovery, Uncategorized

≈ 26 Comments

Tags

ESI Protocols

Periodically, I strive to pen something practical and compendious on electronic evidence and eDiscovery, drilling into a topic, that hasn’t seen prior comprehensive treatment.  I’ve done primers on metadata, forms of production, backup systems, databases, computer forensics, preservation letters, ESI processing, email, digital storage and more, all geared to a Luddite lawyer audience.  I’ve long wanted to write, “The Annotated ESI Protocol.” Finally, it’s done.

The notion behind the The Annotated ESI Protocol goes back 40 years when, as a fledgling personal injury lawyer, I found a book of annotated insurance policies.  What a prize!  Any plaintiff’s lawyer will tell you that success is about more than liability, causation and damages; you’ve got to establish coverage to get paid.  Those annotated insurance policies were worth their weight in gold.

As an homage to that treasured resource, I’ve sought to boil down decades of ESI protocols to a representative iteration and annotate the clauses, explaining the “why” and “how” of each.  I’ve yet to come across a perfect ESI protocol, and I don’t kid myself that I’ve crafted one.  My goal is to offer lawyers who are neither tech-savvy nor e-discovery aficionados a practical, contextual breakdown of a basic ESI protocol–more than simply a form to deploy blindly or an abstract discussion.  I’ve seen thirty-thousand-foot discussions of protocols by other commentators, yet none tied to the document or served up with an ESI protocol anyone can understand and accept. 

It pains me to supply the option of a static image (“TIFF+”) production, but battleships turn slowly, and persuading lawyers long wedded to wasteful ways that they should embrace native production is a tough row to hoe. My intent is that the TIFF+ option in the example sands off the roughest edges of those execrable images; so, if parties aren’t ready to do things the best way, at least we can help them do better.

Fingers crossed you’ll like The Annotated ESI Protocol and put it to work. Your comments here are always valued.

Share this:

  • Email a link to a friend (Opens in new window) Email
  • Print (Opens in new window) Print
  • Share on X (Opens in new window) X
  • Share on Facebook (Opens in new window) Facebook
  • Share on LinkedIn (Opens in new window) LinkedIn
Like Loading...
Follow Ball in your Court on WordPress.com

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 2,241 other subscribers

Recent Posts

  • Fifteen Years in Your Court August 21, 2026
  • The AI Protective Order Double Standard July 27, 2026
  • Drafting RFPs for Robots to Read July 20, 2026
  • A Refresh of the Annotated ESI Protocol May 1, 2026
  • Free at Last: Ditching TurboTax for FreeTaxUSA April 5, 2026

Archives

RSS Feed RSS - Posts

CRAIGBALL.COM

Categories

EDD Blogroll

  • eDiscovery Today (Doug Austin)
  • Basics of E-Discovery (Exterro)
  • Sedona Conference
  • eDiscovery Journal (Greg Buckles)
  • The Relativity Blog
  • Minerva 26 (Kelly Twigger)
  • Illuminating eDiscovery (Lighthouse)
  • Complex Discovery (Rob Robinson)
  • GLTC (Tom O'Connor)
  • Corporate E-Discovery Blog (Zapproved )
  • E-D Team (Ralph Losey)
  • CS DISCO Blog
  • E-Discovery Law Alert (Gibbons)

Admin

  • Create account
  • Log in
  • Entries feed
  • Comments feed
  • WordPress.com

Enter your email address to follow Ball in Your Court and receive notifications of new posts by email.

Website Powered by WordPress.com.

  • Subscribe Subscribed
    • Ball in your Court
    • Join 2,093 other subscribers
    • Already have a WordPress.com account? Log in now.
    • Ball in your Court
    • Subscribe Subscribed
    • Sign up
    • Log in
    • Report this content
    • View site in Reader
    • Manage subscriptions
    • Collapse this bar
Loading Comments...
%d