
Open
Posted
•
Ends in 2 days
Paid on delivery
I’m building an AI-powered service that can read batches of PDFs, Word files, and spreadsheets, then instantly surface the information I ask for. The core functions I need are: • Accurate data extraction from those three file types • Concise English-language summaries generated on demand • Reliable keyword and phrase identification that feeds a searchable index Everything will run in English for now, so no multilingual processing is required. I’m open to whichever stack you prefer—Python, Node, or another language—as long as modern NLP libraries (think PyPDF, docx, pandas, spaCy, transformers, LangChain, etc.) are used cleanly and the code is well-commented. Here’s how I picture the workflow: I upload a document or folder, the system parses each file, stores the structured output in a database or vector store, and exposes a simple interface (CLI, web dashboard, or API endpoint) where I can fire off queries such as “summarize section 3,” “list all invoice totals,” or “show recurring keywords across Q2 reports.” Fast response time and clear, formatted results are key. Deliverables 1. End-to-end working prototype with source code 2. Setup instructions and dependency list 3. Brief usage guide demonstrating extraction, summary, and keyword search on sample files 4. Short video or screenshare walk-through confirming the above features I can provide sample documents for testing. If you’ve built something similar—especially with PDF text extraction quirks or mixed spreadsheet/Word inputs—let me know; that experience will be a big plus. Importnat: I will host applicaytion on my Server. So, sugegstion on appropriate light weight LLM needed.
Project ID: 40568591
67 proposals
Open for bidding
Remote project
Active 57 yrs ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs