Multimodal Financial Document Understanding System
Built a multimodal document-understanding pipeline using open-source vision-language models to extract text, tables, charts, signatures, and layout from multi-page financial documents, replacing traditional OCR-based extraction. Orchestrated the multi-stage pipeline using LangGraph, built a CrewAI-based validation agent using the ReAct pattern to verify extracted fields, and structured outputs into machine-readable JSON.
// INTERVIEW ENGINE