← back to experience

SOFTWARE / AI ENGINEER · 2024-NOW · PHILADELPHIA, PA

Cobbs Creek Healthcare

I worked across five focus areas: automated Databricks pipelines, a data-quality dashboard, an LLM-powered internal chatbot, a dementia aid patient app, and patient-journey modeling. The common thread is making complex healthcare data easier to process, monitor, investigate, and act on.

  • Databricks · PySpark
  • LLM / agentic workflows
  • Vector DB · retrieval
  • Modeling · segmentation

Note: details and visuals on this page are generalized or recreated to protect confidential business and healthcare data.

TECH OVERVIEW

DATA ENGINEERING

  • Databricks
  • PySpark
  • Python
  • SQL
  • Databricks Jobs

AI / LLM

  • Agentic workflows
  • RAG
  • Prompt optimization
  • Summarization

RETRIEVAL

  • Vector databases
  • Structured databases
  • Retrieval flows

MODELING

  • Clustering
  • Logistic regression
  • Anomaly detection
  • Feature engineering

PRODUCT

  • Dashboards
  • Investigation flows
  • UI/UX
  • Reporting

CROSS-PROJECT SKILLS

Data Engineering

Automated workflows that made recurring processing reliable, scalable, auditable.

Applied AI

LLM workflows connecting search, retrieval, summarization & explanation.

Dashboards + Product

Turned complex data outputs into usable investigation flows.

Modeling + Segmentation

Feature engineering, clustering & regression for forecasting and insights.

Healthcare Context

Complex healthcare/pharma workflows kept generalized and reviewable.

FIVE WORK PROJECTS

1. Databricks Data Pipelines

I built and maintained Databricks-based workflows that automated recurring processing, validation, and audit steps across 50+ healthcare/pharma datasets, replacing manual, handoff-heavy processes with repeatable jobs that run consistently, surface issues earlier, and feed downstream reporting and investigation.

50+

datasets automated

Recurring Databricks workflows to process, validate, and audit pharma data sources.

faster recurring runs

Automated repeatable processing and validation, cutting manual handoff steps.

PROBLEM

Manual, handoff-heavy processing made recurring data slow to run and hard to trust or audit.

WHAT I BUILT

Parameterized Databricks Jobs with PySpark/SQL transforms, reusable validation, and audit tables.

RELIABILITY

Job parameters, task values, concurrent for-each runs, and versioned audit tables for transparency.

IMPACT

Earlier issue detection, less manual handoff, and consistent inputs for downstream dashboards.

PIPELINE ARCHITECTURE

  1. Source SystemsInternal data feeds
  2. OrchestrationScheduled jobs
  3. TransformPySpark / SQL
  4. ValidateBusiness + data quality checks
  5. PublishCurated tables
  1. Audit LogsCaptures run status, step status, and versions
  2. Dashboards / ReportsOperational visibility and insights

TECH & SKILLS

  • Databricks
  • PySpark
  • Python
  • SQL
  • Databricks Jobs
  • Audit tables

2. Data Quality Dashboard

I led development of a data-quality dashboard that helped teams review issues, investigate anomalies, and understand where unusual changes were happening. It organizes anomaly outputs into a usable investigation flow, from high-level issue down to the relevant accounts, territories, or aggregation levels.

monitoring coverage

Extended anomaly and DQ checks across more sources and aggregation levels.

Recreated data-quality operations dashboard: open issues, checks reviewed, escalations, data freshness, weekly anomaly volume, severity prioritization, and a national-to-account drill-down. Mock visualization, no real data.

TECH & SKILLS

  • Anomaly detection
  • Data-quality checks
  • Aggregate checks
  • Severity prioritization
  • Investigation drill-down

3. Agentic AI Chatbot

I contributed to an internal agentic AI chatbot that helps users search, retrieve, summarize, and explain information across business workflows. My work focused on specific agents and tools: insurance-policy search, targeted web scraping, summarization, prompt optimization, and retrieval-backed explanation flows.

Agent workflow diagram: a user question goes to a coordinator agent that calls web search, knowledge base, and database/API tools; retrieved context feeds LLM reasoning, summarization, and explanation into a final response, supported by memory, tool results, and a feedback loop.

TECH & SKILLS

  • LLM agents & tools
  • RAG
  • Vector databases
  • Web scraping
  • Prompt optimization
  • Summarization

4. Dementia Aid Patient App

As one of three primary developers on a dementia aid patient app supporting patient and caregiver workflows, I focused on the vector database layer, database structure, and UI/UX, shaping how information is stored, retrieved, and presented in a more accessible way. A prototype/product build, not a clinically validated medical product.

Dementia aid app overview: a patient and caregiver companion app connecting patient profile, memory support, caregiver notes, vector database, database design, and accessible UI/UX. Prototype/product build, not a clinically validated medical product.

Recreated overview shown above. Patient-facing screenshots are intentionally omitted to protect privacy; no real patient data is shown.

TECH & SKILLS

  • Vector database
  • Database design
  • UI/UX
  • Privacy-aware design

5. Patient Journey Modeling

I supported patient-journey modeling using de-identified account-level data to segment behavior, support forecasting, and identify key HCPs. This involved feature engineering, clustering, and logistic regression to understand patterns across accounts for downstream business analysis.

Patient journey modeling diagram: behavior features, volume trends, and engagement signals feed a journey modeling engine built on de-identified account-level data, producing segment profiles, a forecasting outlook, and key account opportunities via feature engineering, model training, validation, and interpretation.

TECH & SKILLS

  • Feature engineering
  • Clustering
  • Logistic regression
  • Segmentation
  • Forecasting support

CLOSING

Across these five areas my focus was building practical systems: pipelines that run reliably, dashboards that support investigation, AI tools that connect retrieval and explanation, and models that make complex healthcare data more understandable.