“I cannot live without looking over my shoulder,” a survivor of sexual abuse told US legislators in May, describing how her personal data was made public in January when the Department of Justice released hundreds of thousands of files tied to financier Jeffrey Epstein. “I can only imagine the lasting repercussions this error will have on my life.”

Three days after the DOJ published the largest of twelve Epstein datasets on January 30, a department attorney informed a federal court that “human or technical error” contributed to the release of personally identifying information of individuals who had been sexually abused by Epstein and his associates. Consequently, the DOJ said it would withdraw, review, and redact flagged documents before republishing them.

DW’s investigative team, together with its Innovation and Data units, spent months examining thousands of documents released by the DOJ. Six months after the release, they discovered dozens of files still containing details—such as names, faces, and email addresses—that could be used to identify survivors, witnesses, informants, and others.

DOJ Epstein Library

By mid‑February, DW had harvested more than 800,000 files from the DOJ’s Epstein Library—a difficult undertaking because the archive was constantly shifting, with new files added, others removed, and altered files reappearing after days of absence.

It was only after DW obtained an earlier archive of the same material hosted by a public data‑preservation project on GitHub that they realized over 500,000 files had been deleted between the initial release and their scrape.

To confirm the integrity of the original release, DW generated a unique digital fingerprint—known as a hash—for every document. Hashes verify that a file matches the original and has not been altered. By comparing these fingerprints with documents harvested from the DOJ site roughly two weeks later, DW could determine which files remained unchanged and which had been removed from the site.

DW also noticed that some files had changed because their hash differed; they retained the same filename but were either larger or smaller than the original. Upon reviewing this subset, DW found that these represented documents that had undergone additional redaction—or, conversely, had information added.

In one example, a survivor recounted how Epstein had abused her for years while she was a minor, in testimony given to lawyers. Her name was redacted throughout the transcript until the final page, where a lawyer thanked her for her testimony and spoke her name aloud; that name remains unredacted in the published version.

Her exposed name was not an isolated incident.

US lawmakers created strict rules on how the Epstein files should be published, such as ensuring that personally identifying information of victims be redacted to protect them from public exposureImage: picture-alliance/dpa/M. Reynolds

Deep audio analysis

To examine audio files—including victim statements and tip‑offs—DW’s investigative unit partnered with the Fraunhofer Institute for Digital Media Technology’s Media Distribution and Security research group in Ilmenau, Germany. The MDS focuses on audio authenticity, manipulation, and speech‑synthesis detection.

After eliminating duplicate files from both the public preservation project and DW’s own scrape, more than 150 audio files remained, which the MDS’s forensic audio analysts examined using cutting‑edge tools developed at the institute. They uncovered digital traces left by audio production and editing processes that are sometimes imperceptible to the human ear.

This work involved pinpointing sections of a recording that had been altered or reused, determining whether a recording was a modified version of another file, and assessing whether audio files might have been captured with the same microphone or recording setup.

Forensic audio researcher Milica Gerhardt noticed something distinctive: redaction styles varied across the files, and the execution of those redactions was inconsistent.

“If you listen to these audio files, it’s clear that a great deal of personal information has not been redacted,” Gerhardt remarked.

Using the Audio Provenance Suite App, Gerhardt identified two files in the same dataset with identical content, but one was redacted and the other unredactedImage: Stefan Czimmek/DW

In one instance, an individual called an official tip‑off hotline to provide information for an investigation into Epstein’s abuse. In the redacted version, the caller’s name is audible while their phone number is bleeped out.

The same dataset, however, also contains an alternative version of the audio file in which the caller’s full name and direct contact details can be heard. Alleged witnesses or informants appear to have been excluded from the protections afforded to victims.

“They could and should have done a better job with the redactions,” said Patrick Aichroth, who leads the MDS group at the Fraunhofer Institute, which has advised German judicial authorities on audio‑related evidence.

“There is a reason such material is handled with extreme care in court proceedings,” Aichroth added. “Once sensitive or identifying information is disclosed, the damage cannot be undone.”

Redactions across audio files were inconsistent and revealed sensitive information, according to audio forensics researchers.Image: Stefan Czimmek/DW

DOJ’s redaction failures

DW’s investigative team presented its findings to the DOJ, including examples where survivors’ personally identifiable information remained in the official Epstein Library.

The DOJ offered no response.

In April, the DOJ’s Office of the Inspector General—an independent watchdog charged with investigating misconduct within the department—announced that it was “initiating an audit of DOJ’s compliance with the Epstein Files Transparency Act.”

“Our preliminary goal is to assess the DOJ’s procedures for identifying, redacting, and releasing records in its possession as required by the Act,” the office stated.

In July, several legislators introduced the Epstein Files Transparency Act II, which would permit sexual‑abuse survivors and the governments of U.S. states to sue the DOJ for failing to obey the rules set out in the original Transparency Act, which had promised to shield survivors’ personally identifying information.

Survivors whose data was released unredacted have since reported harassment, humiliation, and physical threats. Because the disclosed information remains accessible on public platforms—and, in some instances, still resides on the DOJ website—their lives continue to be upended.

“While the wealthy and powerful stay protected by redactions, my name was exposed to the world,” said the sexual‑abuse survivor who testified before Congress in May.

“The evidence is right here,” she said, “yet those in power would rather see us suffer socially, emotionally, and physically than admit their own complicity.”

Edited by: Mathias Bölinger, Milan Gagnon

Technical and data analysis by: Andreas Giefer

Data journalists: Gianna Grün, Rodrigo Menegat Schuinski

Fact checking by: Birgitta Schülke

How the Epstein files have opened a survivor’s deep wounds

To view this video please enable JavaScript, and consider upgrading to a web browser that supports HTML5 video



Source link

Exit mobile version