
In Progress
Posted
Paid on delivery
I’m looking to run a complete multimodal analysis of a body of spoken content that arrives in three different forms: raw audio, corresponding video, and existing text transcripts. My aim is to fuse these streams so we can extract deeper insights than any one modality would reveal on its own. Here’s what I need from you: • Develop or adapt models that synchronise audio, video, and text, producing aligned outputs ready for further study. • Apply speech-to-text (where transcripts are missing or need verification), gauge recognition accuracy, and layer in sentiment or other paralinguistic signals as appropriate. • Package the workflow in reproducible Python notebooks or scripts (PyTorch/TensorFlow, Whisper, Kaldi, or similar libraries are fine), with clear comments and a short read-me explaining how to rerun everything on fresh data. • Deliver a concise analytical report that highlights key findings, method choices, and any recommendations for improving future recordings. Acceptance criteria 1. All input files (audio, video, transcripts) are time-aligned and stored in a structured directory. 2. Accuracy and sentiment scores are presented in tabular form, accompanied by at least one visual dashboard or interactive plot. 3. Code executes end-to-end on my machine with a single command and no missing dependencies. If you have experience blending audio, visual, and textual cues into a unified speech analysis pipeline, I’d love to see how you would tackle this.
Project ID: 40529732
44 proposals
Remote project
Active 4 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs