bulk_extractor – simsong
这是开发代码树。生产版本下载位于:
关键指标一览
README 详细介绍

bulk_extractor is a high-performance digital forensics exploitation
tool. It is a "get evidence" button that rapidly scans any kind of
input (disk images, files, directories of files, etc) and extracts
structured information such as email addresses, credit card numbers,
JPEGs and JSON snippets without parsing the file system or file system
structures. The results are stored in text files that are easily
inspected, searched, or used as inputs for other forensic
processing. bulk_extractor also creates histograms of certain kinds of
features that it finds, such as Google search terms and email
addresses, as previous research has shown that such histograms are
especially useful in investigative and law enforcement applications.
Unlike other digital forensics tools, bulk_extractor probes every byte of data to see if it is the start of a
sequence that can be decompressed or otherwise decoded. If so, the
decoded data are recursively re-examined. As a result, bulk_extractor can find things like BASE64-encoded JPEGs and
compressed JSON objects that traditional carving tools miss.
This source tree builds bulk_extractor 2.2.0. For production use, prefer a tested release from https://github.com/simsong/bulk_extractor/releases.
Building bulk_extractor
=========================
We recommend building from sources. We provide a number of bash scripts in the etc/ directory that will configure a clean virtual machine:
git clone https://github.com/simsong/bulk_extractor.git
./bootstrap.sh
./configure
make
make check
make install
For detailed instructions on installing packages and building bulk_extractor, read the wiki page here:
https://github.com/simsong/bulk_extractor/wiki/Installing-bulk_extractor
For more information on bulk_extractor, visit: https://forensics.wiki/bulk_extractor
Tested Configurations
=====================
This release of bulk_extractor requires C++17. Current validation covers:
- Apple Silicon macOS (local build and
make distcheck, 2026-07-19) - Ubuntu 22.04 and current macOS GitHub Actions runners
Older platform preparation scripts under etc/ are not equivalent to current
CI support and may require maintenance.
Tested Configurations Which bulk_extractor Does Not Work
========================================================
- Debian 10 (not supported for native builds)
RECOMMENDED CITATION
====================
If you are writing a scientific paper and using bulk_extractor, please cite it with:
Garfinkel, Simson, Digital media triage with bulk data analysis and bulk_extractor. Computers and Security 32: 56-72 (2013)
@article{10.5555/2748150.2748581,
author = {Garfinkel, Simson L.},
title = {Digital Media Triage with Bulk Data Analysis and Bulk_extractor},
year = {2013},
issue_date = {February 2013},
publisher = {Elsevier Advanced Technology Publications},
address = {GBR},
volume = {32},
number = {C},
issn = {0167-4048},
journal = {Comput. Secur.},
month = feb,
pages = {56–72},
numpages = {17},
keywords = {Digital forensics, Bulk data analysis, bulk_extractor, Stream-based forensics, Windows hibernation files, Parallelized forensic analysis, Optimistic decompression, Forensic path, Margin, EnCase}
}
ENVIRONMENT VARIABLES
=====================
The following environment variables can be set to change the operation of bulk_extractor:
|Variable|Behavior|
|--------|--------|
|DEBUG_BENCHMARK|Include CPU benchmark information in report.xml file|
|DEBUG_NO_SCANNER_BYPASS|Disables scanner bypass logic that bypasses some scanners if an sbuf contains ngrams or does not have a high distinct character count.|
|DEBUG_HISTOGRAMS|Print debugging information on file-based histograms.|
|DEBUG_HISTOGRAMS_NO_INCREMENTAL|Do not use incremental, memory-based histograms.|
|DEBUG_PRINT_STEPS|Prints to stdout when each scanner is called for each sbuf|
|DEBUG_SCANNER_DUMP_DATA|Hex-dump each sbuf that is to be scanned.|
|DEBUG_SCANNERS_IGNORE|A substring used to identify scanners to ignore. Useful for debugging unit tests.|
Other hints for debugging:
- Run -xall to run without any scanners.
- Run with a random sampling of 0.001% to debug reading image size and a few quick seeks.
LOADABLE SCANNERS
=================
bulk_extractor loads scanner modules named scan_.so (or scan_.dylib on
macOS and scan_*.dll on Windows) from the directories supplied with -P or in
BE_PATH. A module exports this C-linkage factory:
extern "C" scanner_t *bulk_extractor_scanner_v1();
The factory returns a normal scanner_t function. Build modules against the
same bulk_extractor source version as the executable; the scanner's PHASE_INIT
handler must call sp.check_version(). Modules remain loaded until scanner
cleanup completes.
BUILDING ON WINDOWS
===================
Native Windows builds of bulk_extractor are not currently supported.
The Windows MinGW build GitHub Actions workflow cross-compiles the executable
on Ubuntu for every pull request and uploads bulk_extractor64.exe in thebulk_extractor-windows-x86_64 artifact. The workflow verifies that the PE
executable does not import the MinGW, RE2, Abseil, Expat, zlib, or GNU crypto
runtime DLLs. It does not use Wine or run the resulting Windows executable.
The CI artifact is currently built without libewf, so it does not include E01
support.
The traditional repository path is cross-compilation from Fedora; start with a
clean VM and use these commands:
$ git clone https://github.com/simsong/bulk_extractor.git
$ cd bulk_extractor/etc
$ bash CONFIGURE_FEDORA36_win64.bash
$ cd ..
$ make win64
BULK_EXTRACTOR RELEASE NOTES
============================
Release 2.2.0 (July 19, 2026)
Integrated the be20 API and its source dependencies into the bulk_extractor
source tree, eliminating recursive submodule setup. The unified build now
validates bulk_extractor, be20, and DFXML together and fixes thread-pool
shutdown defects found by full-image testing and AddressSanitizer.
Release 2.1.1 (April 26, 2024)
Renamed jpeg_carved feature recorder to jpeg, so that the jpeg carve mode can be set with -S jpeg_carve_mode=2, rather than -S jpeg_carved_carve_mode=2, which was confusing.
Release 2.0
bulk_extractor 2.0 (BE2) is now operational. Although it works with the Java-based viewer, we do not currently have an installer that runs under Windows.
BE2 requires C++17 to compile. The be20 scanner API, dfxml_cpp, utfcpp, and DFXML schema sources are maintained directly in this repository; no recursive submodule checkout is required.
The project took longer than anticipated. In addition to updating to C++17, It was used as an opportunity for massive code refactoring and general increase in code quality, testability and reliability. An article about the experiment will appear in a forthcoming issue of ACM Queue