Spaces:
Sleeping
A newer version of the Gradio SDK is available: 6.22.0
title: Crop Adv
emoji: π
colorFrom: blue
colorTo: yellow
sdk: gradio
sdk_version: 6.8.0
app_file: app.py
pinned: false
license: mit
πΎ Maharashtra Crop Recommendation System
An AI-powered crop advisory tool that recommends the most suitable crops for a Maharashtra farmer based on their soil chemistry, local weather, district & season, water supply, and budget β and explains why each crop is recommended in plain language.
π§ How It Works: The Full Pipeline
Step 1 β Two Models Run in Parallel
When you click Find Best Crops For Me, two independent models are run simultaneously:
Model 1 β Soil & Climate AI (Random Forest)
- Takes 7 numerical inputs: N, P, K, Temperature, Humidity, pH, Rainfall
- Runs them through a pre-trained Random Forest Classifier (
model1_npk.pkl) - Outputs a probability score (0β1) for every crop it knows (22 crops)
- Example:
{"rice": 0.42, "maize": 0.28, "cotton": 0.08, ...} - The label encoder (
model1_label_encoder.pkl) maps numeric class indices β crop names - Training: 2,200 samples Γ 22 crops, 99.6% cross-validation accuracy (5-fold stratified)
Model 2 β Regional History Scoring (Weighted Formula)
- Looks up the selected District + Season combination in
model2_full_scored.csv - Returns a pre-computed suitability score for each crop based on 26 years (1997β2022) of Maharashtra district-level crop production data
- No ML here β this is a transparent, rule-based scoring formula:
Suitability Score =
0.35 Γ Crop Frequency % (how often this crop is grown here)
+ 0.25 Γ Historically Grown Flag (grown in β₯50% of years = 1, else 0)
+ 0.20 Γ Normalised Avg Yield (yield relative to best crop in district)
+ 0.10 Γ Dominance Rate (share of total district area)
+ 0.05 Γ Recent Years Grown (grown post-2015 = recency bonus)
+ 0.05 Γ Yield Stability (lower variance = higher score)
- Covers 35 districts Γ 4 seasons Γ 33 crops = 1,603 scored combinations
Step 2 β Scores Are Combined
Final Score = 0.60 Γ Model1_probability + 0.40 Γ Model2_suitability_score
- Model 2 scores are normalised to 0β1 before combining (divide by max score in that district/season)
- Crops are then ranked highest β lowest by Final Score
- Only crops present in
CROP_META(hardcoded dictionary, see below) are kept
Step 3 β Budget & Land Filtering
For each ranked crop:
| Condition | Result |
|---|---|
input_cost Γ land_ha β€ budget |
budget_status = "within" β shown fully |
input_cost Γ land_ha > budget |
affordable_ha = budget / input_cost β shown with partial land warning |
affordable_ha < 0.1 ha |
Crop is skipped entirely |
Step 4 β Risk Scoring (4-Dimensional)
Every crop gets a risk score from 0β100 across four dimensions, then averaged equally:
| Dimension | Weight | How It's Calculated |
|---|---|---|
| Weather Risk | 25% | Penalises extreme temp (<8Β°C or >42Β°C), extreme humidity, and rainfall mismatches vs crop water need |
| Market Risk | 25% | Based on hardcoded demand: High=20, Medium=40, Low=65 |
| Budget Risk | 25% | Ratio of total_cost / budget_max mapped to risk bands |
| Water Risk | 25% | Gap between crop's irrigation need vs farmer's available supply |
Irrigation need levels (hardcoded per crop):
"high"β Rice, Sugarcane, Banana, Grapes, Jute, Apple"medium"β Cotton, Maize, Wheat, Onion, Tomato, etc."low"β Chickpea, Bajra, Jowar, Pulses, Oilseeds, etc.
Irrigation supply levels (farmer input):
- Assured (Canal / River) = 3
- Borewell / Pump = 2
- Seasonal / Rain-fed = 1
- No Irrigation (Dryland) = 0
Risk label thresholds:
- Score < 35 β π’ Low
- Score 35β64 β π‘ Medium
- Score β₯ 65 β π΄ High
Step 5 β Confidence Labelling
Each result is tagged with an AI confidence level based on Model 1's raw probability:
| Model 1 Score | Label | Meaning |
|---|---|---|
| β₯ 35% | High confidence | Soil & climate strongly match |
| 12β34% | Moderate confidence | Partial soil match, regional history confirms |
| < 12% | Based on regional history | Driven primarily by Model 2 |
Step 6 β "Why We Recommend This" Explanation
Every card generates 5 plain-language bullet points dynamically from real data β nothing is templated or hardcoded:
| Bullet | Data Source |
|---|---|
| Soil fit | m1_score from Model 1 probability |
| Season fit | season_match boolean (best_season == selected season) |
| Water fit | Irrigation gap calculation (need β available) |
| Risk summary | risk label + highest risk_breakdown dimension |
π₯ Inputs
Location & Season
| Input | Type | Options |
|---|---|---|
| District | Dropdown | 35 Maharashtra districts |
| Season | Dropdown | Kharif (JunβSep), Rabi (OctβMar), Summer (AprβJun), Whole Year |
Farm Details
| Input | Type | Range |
|---|---|---|
| Land (Hectares) | Number | 0.1 β 100 |
| Total Budget (βΉ) | Number | βΉ1,000+ |
| Water source | Radio | Assured / Borewell / Seasonal / Dryland |
Soil Test Results (from Krishi Vigyan Kendra or soil lab)
| Input | Range | Unit |
|---|---|---|
| Nitrogen (N) | 0 β 300 | mg/kg |
| Phosphorus (P) | 0 β 300 | mg/kg |
| Potassium (K) | 0 β 300 | mg/kg |
| Soil pH | 3.5 β 9.5 | β |
Local Weather
| Input | Range | Unit |
|---|---|---|
| Avg Temperature | 5 β 50 | Β°C |
| Avg Humidity | 10 β 100 | % |
| Annual Rainfall | 20 β 3000 | mm/year |
π€ Output: What Each Card Shows
| Element | Source | Notes |
|---|---|---|
| Crop name + medal | Combined score ranking | π₯π₯π₯ for top 3 |
| Best season | CROP_META hardcoded |
e.g. "Kharif", "Rabi", "Whole Year" |
| Days to harvest | CROP_META hardcoded |
Approximate growing period |
| Budget badge | Budget filter logic | β Fits / β οΈ Partial |
| Risk badge | Risk score calculation | Low / Medium / High |
| "Why we recommend this" | Live data from both models | 4 dynamic bullet points |
| Irrigation verdict | Irrigation gap formula | β / β οΈ / β |
| Cost overview | CROP_META Γ land_ha |
Input cost per ha and total |
| Market demand | CROP_META hardcoded |
High / Medium / Low |
| Risk breakdown bars | 4-dimensional risk formula | 0β100 per dimension |
| Model confidence | Model 1 probability | High / Moderate / Regional |
| Score breakdown | Both model scores | Soil %, Regional %, Combined % |
ποΈ What Is Hardcoded (CROP_META)
The following per-crop values are static estimates based on Maharashtra Agriculture Department averages. They are not fetched from any live source:
| Field | Description | Example (Rice) |
|---|---|---|
input_cost |
Estimated cost to grow 1 hectare (βΉ) | βΉ30,000 |
yield_qtl |
Expected yield in quintals per hectare | 25 qtl/ha |
demand |
Market demand category | "High" |
days |
Days from sowing to harvest | 120 days |
best_season |
Ideal growing season | "Kharif" |
Full crop list in CROP_META (50 crops): Rice, Maize, Chickpea, Kidneybeans, Pigeonpeas, Mothbeans, Mungbean, Blackgram, Lentil, Pomegranate, Banana, Mango, Grapes, Watermelon, Muskmelon, Apple, Orange, Papaya, Coconut, Cotton, Jute, Coffee, Arhar/Tur, Bajra, Castor Seed, Gram, Groundnut, Jowar, Linseed, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds, Cotton(Lint)
β οΈ Note:
input_cost,yield_qtl, anddemandare estimates. Actual values vary by year, variety, and farming practices.
ποΈ Files in This Space
| File | Description |
|---|---|
app.py |
Main Gradio application β all logic lives here |
model1_npk.pkl |
Trained Random Forest classifier (soil & climate, 22 crops) |
model1_label_encoder.pkl |
Label encoder β maps class indices β crop names |
model2_full_scored.csv |
Pre-computed regional suitability scores (35 districts Γ 4 seasons Γ 33 crops) |
requirements.txt |
Python dependencies |
README.md |
This file |
π¦ Requirements
gradio>=4.0.0
joblib
numpy
pandas
scikit-learn
β οΈ
scikit-learnmust be listed even though it is never directly imported βjoblibneeds it to deserialise the.pklmodel files.
ποΈ System Architecture
Farmer Inputs (soil + weather + location + farm details)
β
ββββΊ Model 1 (Random Forest)
β model1_npk.pkl + model1_label_encoder.pkl
β β probability score per crop (22 crops)
β
ββββΊ Model 2 (Regional Scoring Table)
β model2_full_scored.csv
β β suitability score per crop (33 crops, filtered by district + season)
β
ββββΊ Combined Score (60% Model 1 + 40% Model 2)
β
βΌ
Ranked crop list
β
ββββΊ Budget filter (skip if < 0.1 ha affordable)
ββββΊ Risk scoring (weather + market + budget + water)
ββββΊ Confidence labelling (based on Model 1 score)
ββββΊ "Why" explanation (dynamically generated from real data)
β
βΌ
HTML output cards (mobile-first UI)
π± Crops by Model Coverage
Model 1 β Soil/Climate (22 crops): Apple, Banana, Blackgram, Chickpea, Coconut, Coffee, Cotton, Grapes, Jute, Kidneybeans, Lentil, Maize, Mango, Mothbeans, Mungbean, Muskmelon, Orange, Papaya, Pigeonpeas, Pomegranate, Rice, Watermelon
Model 2 β Regional (33 crops): Arhar/Tur, Bajra, Banana, Castor Seed, Cotton, Gram, Grapes, Groundnut, Jowar, Linseed, Maize, Mango, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Rice, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds
Crops in both models get a fully blended score. Crops in only one model get a 0 from the missing model but are still shown if the other model scores them highly.
π Data Coverage
| Details | |
|---|---|
| State | Maharashtra, India |
| Districts | 35 |
| Seasons | Kharif (JunβSep), Rabi (OctβMar), Summer (AprβJun), Whole Year |
| Historical data range | 1997 β 2022 (26 years) |
| Regional data source | Maharashtra district-level crop production statistics |
| Soil model training data | 2,200 samples across 22 crops (100 per crop) |
| Soil model accuracy | 99.6% cross-validation (5-fold stratified) |
π·οΈ Suitability & Risk Labels
Regional Suitability (Model 2 raw score)
| Label | Score Range | Meaning |
|---|---|---|
| π’ Highly Suitable | β₯ 75% | Strong regional farming history |
| π‘ Moderately Suitable | 50β74% | Good candidate, moderate history |
| π Low Suitability | 30β49% | Possible but limited evidence |
| π΄ Not Recommended | < 30% | Little or no regional history |
Overall Risk (combined score)
| Label | Score Range | Action |
|---|---|---|
| π’ Low | < 35 | Safe choice for this season |
| π‘ Medium | 35β64 | Proceed with planning for main risk factor |
| π΄ High | β₯ 65 | Consider reducing land allocation |
π¨βπ» Built For
Maharashtra farmers and agricultural advisors. Designed to be used on mobile phones in the field β large text, simple language, clear YES/NO verdicts, with detailed breakdown available on demand.
Hackathon criteria addressed:
- β Integration of multiple data inputs (soil, weather, water, location, budget)
- β Crop suitability prediction via dual-model AI pipeline
- β Risk-aware scoring with 4-dimensional uncertainty handling
- β Clear, simple recommendation outputs with plain-language explanations
- β Farmer-friendly, mobile-responsive interface
- β Scalable architecture (local model files, no external API dependencies)