A web application for Indonesian coffee farmers. It predicts coffee pH and quality from the climate, topography, and soil data of their own plantations.
Coffee farmers in Indonesia rarely have access to laboratory testing. The character of their coffee, including its acidity, only becomes clear once the harvest reaches a buyer and is judged by cupping. Yet what shapes that character was present long before, out in the plantation: altitude, daily temperature, rainfall, slope, and soil condition. That information is actually available through public APIs, but it arrives as raw numbers that mean nothing to a farmer.
I built Prediksi Kopi so farmers can see an early estimate of their coffee's character using data from their own plantation, without having to gather those numbers one by one.
The application runs as two separate services. The web application is built with Laravel and stores plantations, prediction history, the Library content, and the forum. The second is a machine learning service in FastAPI that loads the scikit-learn models and can only be called by the web application with an API key, never directly from the browser.
When a farmer adds a plantation, they pick a location down to village level from the Kemendagri region data I imported into the database (91,599 rows), then drag a map marker to the plantation itself. Once saved, three groups of environmental data are fetched in the background through a queue: climate from the Open-Meteo archive, altitude from Open-Meteo Elevation, and soil condition from SoilGrids. Each group records its source (automatic or manual) along with the time it was fetched, so the farmer knows which figures came from an API and which they entered themselves.
The biggest obstacle was not the code. The real coffee pH data I had amounted to four points from regional laboratory tests. That is nowhere near enough to train a model. I could have trained something anyway and called it a pH prediction, but the number would have misled farmers. So I took another route: the machine learning models are trained on quality score and sensory acidity score using the CQI dataset (1,338 rows), while the pH value is shown as an estimate from a formula and openly labelled as not being a model output. On the prediction detail page I added a field for lab-tested pH, so real data can accumulate slowly. A pH model will only be trained once there is enough of it.
I also show the model metrics as they are. The quality model scores an R2 of 0.16 and an RMSE of 2.81 on the test set, the acidity model an R2 of 0.17 and an RMSE of 0.29. Those numbers are indeed low, and on the results page I describe them as a rough estimate rather than a certainty.
The next obstacle came from the environmental APIs. Their units are not directly usable. Slope from the API is in degrees while field extension work uses percentages, and the sand, silt, and clay composition from SoilGrids arrives in grams per kilogram, which has to be converted before mapping onto the 12 USDA texture classes in Indonesian. A division mistake in this part once made every soil read as the same class.
While seeding the initial production data, I hit HTTP 429 because the ML service limits requests to 60 per minute, and the prediction history is genuinely computed through the model. I added a 1.2 second gap per prediction in the seeder, and the data was stored intact.
The admin panel uses a Bootstrap template that reloads every asset on each page change. I added Turbo Drive so assets load once at the start, then marked the vendor scripts so they are not re-evaluated.
I built this project alone: designing the database schema, the Laravel backend, the interface, the environmental API integration, training both scikit-learn models, the FastAPI service, writing 125 automated tests, and the Dockerfile and deployment to Coolify.
The application runs at prediksikopi.com with two containers (web and ML) and a MySQL database on a single VPS. All automated tests pass before every deployment.
What I want to continue: collecting lab-tested pH from users until there is enough to train a real pH model, then re-evaluating the quality model by adding plantation environment features to it. Only after that does it make sense to talk about accuracy more seriously.