Project 004 · Software / AI · Seminole Innovators

Football
Analytics
AI

A predictive football analytics model built using data preprocessing, feature selection, and machine learning — combined with LLM prompt engineering to accelerate development. Improved model accuracy by 10% and saved over 2 months of engineering time.

Timeline

Jan 2025 – Jan 2026

Context

Seminole Innovators — Board Position

Tools

Python, ML libraries, LLM prompt engineering

10%
Accuracy improvement via feature selection
2+
Months of dev time saved via LLM prompting
ML
Predictive modelling pipeline
LLM
Prompt engineering for code generation

The Problem

Noisy Data Was Killing Model Performance

Raw football datasets contain hundreds of variables — most of which add noise rather than signal. The challenge was identifying which features actually predicted outcomes and building a model that was accurate, not just large.

A secondary challenge was development speed: building, testing, and iterating a machine learning pipeline from scratch with a small team is slow. Finding ways to accelerate without sacrificing quality was equally important.

The Approach

Feature Selection + LLM Accelerated Development

Two parallel strategies drove the results:

  • Data preprocessing & feature selection: Systematically reduced the feature space by identifying which variables had meaningful predictive power ( Touchdowns, yardage gained, quarter) — removing noise that was degrading model accuracy
  • LLM prompt engineering: Used large language models as a code generation accelerator and guide — writing precise prompts to generate, debug, and refactor pipeline components instead of coding everything from scratch

Technical Details

What Was Built

  • Data ingestion and cleaning pipeline to handle raw football statistics
  • Feature selection analysis to narrow high-dimensional data to effective predictors
  • Predictive model training and validation with accuracy benchmarking
  • Prompt engineering workflow for LLM-assisted code generation and debugging

The combination of better features and faster iteration cycles compounded — cleaner data made the model more accurate, and faster development meant more iterations in the same time budget.

Why This Matters for Engineering

Data Skills Transfer Directly to Hardware

The same skills used here — data preprocessing, feature selection, signal-vs-noise thinking — are core to data acquisition systems in aerospace and energy applications. Sensor data from high pressure systems or flight test hardware has the same fundamental problem: large, noisy datasets where identifying meaningful signal is the hard part.

LLM-accelerated development is rapidly becoming a baseline skill in engineering. The ability to write precise prompts that generate usable code reduces iteration time in any technical discipline.

Results & Takeaways

Outcome

Model accuracy improved by 10% through targeted feature selection, and LLM-assisted development saved over 2 months of Development time.

What I Learned

More data and more features don't produce better models — they produce slower, noisier ones. The real work is understanding which inputs actually matter, which is a discipline that applies equally to experimental sensor setups and production data pipelines.

← Previous: Stirling Engine All Projects →