SOFTWARE / AI ENGINEER · 2024-NOW · PHILADELPHIA, PA
Cobbs Creek Healthcare
I worked across five focus areas: automated Databricks pipelines, a data-quality dashboard, an LLM-powered internal chatbot, a dementia aid patient app, and patient-journey modeling. The common thread is making complex healthcare data easier to process, monitor, investigate, and act on.
- Databricks · PySpark
- LLM / agentic workflows
- Vector DB · retrieval
- Modeling · segmentation
Note: details and visuals on this page are generalized or recreated to protect confidential business and healthcare data.
TECH OVERVIEW
DATA ENGINEERING
- Databricks
- PySpark
- Python
- SQL
- Databricks Jobs
AI / LLM
- Agentic workflows
- RAG
- Prompt optimization
- Summarization
RETRIEVAL
- Vector databases
- Structured databases
- Retrieval flows
MODELING
- Clustering
- Logistic regression
- Anomaly detection
- Feature engineering
PRODUCT
- Dashboards
- Investigation flows
- UI/UX
- Reporting
CROSS-PROJECT SKILLS
Data Engineering
Automated workflows that made recurring processing reliable, scalable, auditable.
Applied AI
LLM workflows connecting search, retrieval, summarization & explanation.
Dashboards + Product
Turned complex data outputs into usable investigation flows.
Modeling + Segmentation
Feature engineering, clustering & regression for forecasting and insights.
Healthcare Context
Complex healthcare/pharma workflows kept generalized and reviewable.
FIVE WORK PROJECTS
Databricks Data Pipelines
Automated recurring processing, validation & audit across 50+ datasets.
Data Quality Dashboard
Led development of an anomaly review & investigation workflow.
Agentic AI Chatbot
Agents & tools for search, retrieval, summarization & explanation.
Dementia Aid Patient App
One of three primary devs: vector DB, DB design & UI/UX.
Patient Journey Modeling
Account-level segmentation & forecasting support with clustering + logistic regression.
1. Databricks Data Pipelines
I built and maintained Databricks-based workflows that automated recurring processing, validation, and audit steps across 50+ healthcare/pharma datasets, replacing manual, handoff-heavy processes with repeatable jobs that run consistently, surface issues earlier, and feed downstream reporting and investigation.
50+
datasets automated
Recurring Databricks workflows to process, validate, and audit pharma data sources.
5×
faster recurring runs
Automated repeatable processing and validation, cutting manual handoff steps.
PROBLEM
Manual, handoff-heavy processing made recurring data slow to run and hard to trust or audit.
WHAT I BUILT
Parameterized Databricks Jobs with PySpark/SQL transforms, reusable validation, and audit tables.
RELIABILITY
Job parameters, task values, concurrent for-each runs, and versioned audit tables for transparency.
IMPACT
Earlier issue detection, less manual handoff, and consistent inputs for downstream dashboards.
PIPELINE ARCHITECTURE
- Source SystemsInternal data feeds
- OrchestrationScheduled jobs
- TransformPySpark / SQL
- ValidateBusiness + data quality checks
- PublishCurated tables
- Audit LogsCaptures run status, step status, and versions
- Dashboards / ReportsOperational visibility and insights
TECH & SKILLS
- Databricks
- PySpark
- Python
- SQL
- Databricks Jobs
- Audit tables
2. Data Quality Dashboard
I led development of a data-quality dashboard that helped teams review issues, investigate anomalies, and understand where unusual changes were happening. It organizes anomaly outputs into a usable investigation flow, from high-level issue down to the relevant accounts, territories, or aggregation levels.
4×
monitoring coverage
Extended anomaly and DQ checks across more sources and aggregation levels.

TECH & SKILLS
- Anomaly detection
- Data-quality checks
- Aggregate checks
- Severity prioritization
- Investigation drill-down
3. Agentic AI Chatbot
I contributed to an internal agentic AI chatbot that helps users search, retrieve, summarize, and explain information across business workflows. My work focused on specific agents and tools: insurance-policy search, targeted web scraping, summarization, prompt optimization, and retrieval-backed explanation flows.

TECH & SKILLS
- LLM agents & tools
- RAG
- Vector databases
- Web scraping
- Prompt optimization
- Summarization
4. Dementia Aid Patient App
As one of three primary developers on a dementia aid patient app supporting patient and caregiver workflows, I focused on the vector database layer, database structure, and UI/UX, shaping how information is stored, retrieved, and presented in a more accessible way. A prototype/product build, not a clinically validated medical product.

Recreated overview shown above. Patient-facing screenshots are intentionally omitted to protect privacy; no real patient data is shown.
TECH & SKILLS
- Vector database
- Database design
- UI/UX
- Privacy-aware design
5. Patient Journey Modeling
I supported patient-journey modeling using de-identified account-level data to segment behavior, support forecasting, and identify key HCPs. This involved feature engineering, clustering, and logistic regression to understand patterns across accounts for downstream business analysis.

TECH & SKILLS
- Feature engineering
- Clustering
- Logistic regression
- Segmentation
- Forecasting support
CLOSING
Across these five areas my focus was building practical systems: pipelines that run reliably, dashboards that support investigation, AI tools that connect retrieval and explanation, and models that make complex healthcare data more understandable.