mirror of
https://github.com/sooox-cc/breach-browser.git
synced 2026-10-02 18:08:35 +02:00
A high-performance full-text search engine with a modern web interface for searching large text datasets. Features advanced search syntax, and efficient file chunking.
- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .gitignore | ||
| app.py | ||
| LICENSE | ||
| preview.png | ||
| README.md | ||
| requirements.txt | ||
Breach Browser
A fast and efficient full-text search engine for large text datasets with a clean web interface. Built with Python, Elasticsearch, and Flask.
Features
- 🔍 Full-text search with advanced query syntax
- ⚡ Efficient chunking of large files
- 📊 Detailed search results with file metadata
- 🔧 Progressive loading for faster search results
- 💾 Resume-capable indexing with progress tracking
- 📝 Search highlighting with context
- 📈 Real-time indexing statistics
Requirements
- Python 3.8+
- Docker
- Sufficient storage space for your dataset and indices
- RAM, a whole lot of RAM
Installation
- Clone the repository:
git clone https://github.com/sooox-cc/breach-browser.git
cd breach-browser
- Create and activate a virtual environment:
python -m venv venv
source venv/bin/activate # On Windows: .\venv\Scripts\activate
- Install the required packages:
pip install -r requirements.txt
- Setup Elasticsearch:
Create data Directory:
mkdir -p /opt/elasticsearch/data
chown -R 1000:1000 /opt/elasticsearch/data
Run Elasticsearch container:
docker run -d \
--name elasticsearch \
-p 9200:9200 \
-p 9300:9300 \
-e "discovery.type=single-node" \
-e "ES_JAVA_OPTS=-Xms2g -Xmx2g" \
-e "xpack.security.enabled=false" \
-v /opt/elasticsearch/data:/usr/share/elasticsearch/data \
--ulimit nofile=65535:65535 \
--ulimit memlock=-1:-1 \
elasticsearch:7.17.9
Usage
- Create a directory named
dataand place your text files there:
mkdir -p data
# Add your text files to data/
- Run the application:
python app.py
- Choose whether to reindex your data when prompted
- Access the web interface at
http://localhost:5000
Search Syntax
The search interface supports various query types:
term1 AND term2- Find documents containing both terms"exact phrase"- Find exact phrase matchestest*- Find words starting with "test"term~- Find similar terms (fuzzy search)term1 OR term2- Find documents with either termNOT term- Exclude documents with this term
Configuration
Edit in app.py:
chunk_size: Text chunk size (default: 500KB)batch_size: Index batch size (default: 5)port: Web interface port (default: 5000)
File Structure
breach-browser/
├── app.py # Main application file
├── requirements.txt # Python dependencies
├── static/ # Static assets
│ └── style.css # CSS styles
├── templates/ # HTML templates
│ └── index.html # Main search interface
└── data/ # Directory for text files (not included)
Configuration
The application can be configured by modifying the following parameters in app.py:
chunk_size: Size of text chunks (default: 500KB)batch_size: Number of documents per indexing batch (default: 5)port: Web interface port (default: 5000)index_name: Elasticsearch index name (default: text_documents)data_dir: The path of the data directory (default:/data)
Security Considerations
This tool is intended for local use. When deploying:
- Enable Elasticsearch security features
- Configure proper authentication
- Use HTTPS for the web interface
- Restrict network access appropriately
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Disclaimer
This tool is for educational and research purposes only. Users are responsible for ensuring they have appropriate permissions for any data they index and search.
