Skip to content

ThisIsNotANamepng/search_engine

Repository files navigation

🔍 Stultus - The Stupid Search Engine

Stultus is latin for "stupid"

Stultus is a full search engine system made by UWEC students with hand-built internet crawler, tokenization, and search algorithm. This all started as a project to learn about web crawling, tokenizing, databases, and server hardware.

And to build something cool. 🔥

AI Usage

We are students, the goal of this project is to learn. To that end, we don't use AI to code the scraper. We can use AI to come up with ideas, to guide how we design our code, but we don't use AI to vibe code.

The supporting servers such as the web text storage server and proxy server are exceptions to this, as learning how to implement those efficiently would have limited us to

  • Scraper: Brainstorming only
  • Blocklist serving and parsing: No AI
  • Web text storage server: Major AI
  • Proxy server: Full AI
  • Dashboard: Moderate AI
  • Searching: No AI
  • Docker images: Brainstorming
  • Coordination server: Major AI
  • Docs: No AI

This roughly equates to an average of ~2.4 on a scale of 0-5 of AI usage (not considering size of each task)

Real Good AI Use Rating of 2: Brainstorming Only

Indexes to index

Development

Setup Development Environment

# Linux
git clone https://github.com/ThisIsNotANamepng/search_engine.git
cd search_engine
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
DATABASE_URL="postgres://postgres:postgressPassword@postgressIP:5432/search_engine" python3 fileNameHere.py
# Windows
# you can't scrape on Windows, the scraper uses a function only available on Linux
git clone https://github.com/ThisIsNotANamepng/search_engine.git
cd search_engine
python -m venv venv
venv\Scripts\Activate
pip install -r requirements.txt
$env:DATABASE_URL = "postgres://postgres:postgressPassword@postgressIP:5432/search_engine" 

Project Layout

app.py             Dashboard & searching pages
search.py          Searching functions & algorithms
scrape.py          Main scraping script (relies on utils.py)
utils.py           Functions called to scrape (relies tokenizer)
tokenizer.py       Tokenize the scraped data to be added to the database
seed_urls.csv      Website URLs to jumpstart the crawler queue
templates          
├─ index.html       Main searching interface
├─ search.html      Search results list view
├─ dashboard.html   Statistics list
└─ creators.html    About the creators page
static
└─ ...              All static items that templates rely on.
docs
└─ ...              General documentation things 

Contributing

The main way to contribute is finding a task in the issue tracker or looking through them in the "Project" tab of the repo and opening a pull request.

Fork the repository, make changes, and open a pull request detailing the changes you've made, and we'll work with you to integrate it!

About

Stultus, a search engine for learning the basics of computer science principles

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages