crop-adv / README.md
rishiphadale's picture
Update README.md
89afa7b verified
|
Raw
History Blame Contribute Delete
12.3 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: Crop Adv
emoji: 🌍
colorFrom: blue
colorTo: yellow
sdk: gradio
sdk_version: 6.8.0
app_file: app.py
pinned: false
license: mit

🌾 Maharashtra Crop Recommendation System

An AI-powered crop advisory tool that recommends the most suitable crops for a Maharashtra farmer based on their soil chemistry, local weather, district & season, water supply, and budget β€” and explains why each crop is recommended in plain language.


🧠 How It Works: The Full Pipeline

Step 1 β€” Two Models Run in Parallel

When you click Find Best Crops For Me, two independent models are run simultaneously:

Model 1 β€” Soil & Climate AI (Random Forest)

  • Takes 7 numerical inputs: N, P, K, Temperature, Humidity, pH, Rainfall
  • Runs them through a pre-trained Random Forest Classifier (model1_npk.pkl)
  • Outputs a probability score (0–1) for every crop it knows (22 crops)
  • Example: {"rice": 0.42, "maize": 0.28, "cotton": 0.08, ...}
  • The label encoder (model1_label_encoder.pkl) maps numeric class indices β†’ crop names
  • Training: 2,200 samples Γ— 22 crops, 99.6% cross-validation accuracy (5-fold stratified)

Model 2 β€” Regional History Scoring (Weighted Formula)

  • Looks up the selected District + Season combination in model2_full_scored.csv
  • Returns a pre-computed suitability score for each crop based on 26 years (1997–2022) of Maharashtra district-level crop production data
  • No ML here β€” this is a transparent, rule-based scoring formula:
Suitability Score =
  0.35 Γ— Crop Frequency %          (how often this crop is grown here)
+ 0.25 Γ— Historically Grown Flag   (grown in β‰₯50% of years = 1, else 0)
+ 0.20 Γ— Normalised Avg Yield      (yield relative to best crop in district)
+ 0.10 Γ— Dominance Rate            (share of total district area)
+ 0.05 Γ— Recent Years Grown        (grown post-2015 = recency bonus)
+ 0.05 Γ— Yield Stability           (lower variance = higher score)
  • Covers 35 districts Γ— 4 seasons Γ— 33 crops = 1,603 scored combinations

Step 2 β€” Scores Are Combined

Final Score = 0.60 Γ— Model1_probability + 0.40 Γ— Model2_suitability_score
  • Model 2 scores are normalised to 0–1 before combining (divide by max score in that district/season)
  • Crops are then ranked highest β†’ lowest by Final Score
  • Only crops present in CROP_META (hardcoded dictionary, see below) are kept

Step 3 β€” Budget & Land Filtering

For each ranked crop:

Condition Result
input_cost Γ— land_ha ≀ budget budget_status = "within" β€” shown fully
input_cost Γ— land_ha > budget affordable_ha = budget / input_cost β€” shown with partial land warning
affordable_ha < 0.1 ha Crop is skipped entirely

Step 4 β€” Risk Scoring (4-Dimensional)

Every crop gets a risk score from 0–100 across four dimensions, then averaged equally:

Dimension Weight How It's Calculated
Weather Risk 25% Penalises extreme temp (<8Β°C or >42Β°C), extreme humidity, and rainfall mismatches vs crop water need
Market Risk 25% Based on hardcoded demand: High=20, Medium=40, Low=65
Budget Risk 25% Ratio of total_cost / budget_max mapped to risk bands
Water Risk 25% Gap between crop's irrigation need vs farmer's available supply

Irrigation need levels (hardcoded per crop):

  • "high" β€” Rice, Sugarcane, Banana, Grapes, Jute, Apple
  • "medium" β€” Cotton, Maize, Wheat, Onion, Tomato, etc.
  • "low" β€” Chickpea, Bajra, Jowar, Pulses, Oilseeds, etc.

Irrigation supply levels (farmer input):

  • Assured (Canal / River) = 3
  • Borewell / Pump = 2
  • Seasonal / Rain-fed = 1
  • No Irrigation (Dryland) = 0

Risk label thresholds:

  • Score < 35 β†’ 🟒 Low
  • Score 35–64 β†’ 🟑 Medium
  • Score β‰₯ 65 β†’ πŸ”΄ High

Step 5 β€” Confidence Labelling

Each result is tagged with an AI confidence level based on Model 1's raw probability:

Model 1 Score Label Meaning
β‰₯ 35% High confidence Soil & climate strongly match
12–34% Moderate confidence Partial soil match, regional history confirms
< 12% Based on regional history Driven primarily by Model 2

Step 6 β€” "Why We Recommend This" Explanation

Every card generates 5 plain-language bullet points dynamically from real data β€” nothing is templated or hardcoded:

Bullet Data Source
Soil fit m1_score from Model 1 probability
Season fit season_match boolean (best_season == selected season)
Water fit Irrigation gap calculation (need βˆ’ available)
Risk summary risk label + highest risk_breakdown dimension

πŸ“₯ Inputs

Location & Season

Input Type Options
District Dropdown 35 Maharashtra districts
Season Dropdown Kharif (Jun–Sep), Rabi (Oct–Mar), Summer (Apr–Jun), Whole Year

Farm Details

Input Type Range
Land (Hectares) Number 0.1 – 100
Total Budget (β‚Ή) Number β‚Ή1,000+
Water source Radio Assured / Borewell / Seasonal / Dryland

Soil Test Results (from Krishi Vigyan Kendra or soil lab)

Input Range Unit
Nitrogen (N) 0 – 300 mg/kg
Phosphorus (P) 0 – 300 mg/kg
Potassium (K) 0 – 300 mg/kg
Soil pH 3.5 – 9.5 β€”

Local Weather

Input Range Unit
Avg Temperature 5 – 50 Β°C
Avg Humidity 10 – 100 %
Annual Rainfall 20 – 3000 mm/year

πŸ“€ Output: What Each Card Shows

Element Source Notes
Crop name + medal Combined score ranking πŸ₯‡πŸ₯ˆπŸ₯‰ for top 3
Best season CROP_META hardcoded e.g. "Kharif", "Rabi", "Whole Year"
Days to harvest CROP_META hardcoded Approximate growing period
Budget badge Budget filter logic βœ… Fits / ⚠️ Partial
Risk badge Risk score calculation Low / Medium / High
"Why we recommend this" Live data from both models 4 dynamic bullet points
Irrigation verdict Irrigation gap formula βœ… / ⚠️ / ❌
Cost overview CROP_META Γ— land_ha Input cost per ha and total
Market demand CROP_META hardcoded High / Medium / Low
Risk breakdown bars 4-dimensional risk formula 0–100 per dimension
Model confidence Model 1 probability High / Moderate / Regional
Score breakdown Both model scores Soil %, Regional %, Combined %

πŸ—„οΈ What Is Hardcoded (CROP_META)

The following per-crop values are static estimates based on Maharashtra Agriculture Department averages. They are not fetched from any live source:

Field Description Example (Rice)
input_cost Estimated cost to grow 1 hectare (β‚Ή) β‚Ή30,000
yield_qtl Expected yield in quintals per hectare 25 qtl/ha
demand Market demand category "High"
days Days from sowing to harvest 120 days
best_season Ideal growing season "Kharif"

Full crop list in CROP_META (50 crops): Rice, Maize, Chickpea, Kidneybeans, Pigeonpeas, Mothbeans, Mungbean, Blackgram, Lentil, Pomegranate, Banana, Mango, Grapes, Watermelon, Muskmelon, Apple, Orange, Papaya, Coconut, Cotton, Jute, Coffee, Arhar/Tur, Bajra, Castor Seed, Gram, Groundnut, Jowar, Linseed, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds, Cotton(Lint)

⚠️ Note: input_cost, yield_qtl, and demand are estimates. Actual values vary by year, variety, and farming practices.


πŸ—‚οΈ Files in This Space

File Description
app.py Main Gradio application β€” all logic lives here
model1_npk.pkl Trained Random Forest classifier (soil & climate, 22 crops)
model1_label_encoder.pkl Label encoder β€” maps class indices β†’ crop names
model2_full_scored.csv Pre-computed regional suitability scores (35 districts Γ— 4 seasons Γ— 33 crops)
requirements.txt Python dependencies
README.md This file

πŸ“¦ Requirements

gradio>=4.0.0
joblib
numpy
pandas
scikit-learn

⚠️ scikit-learn must be listed even though it is never directly imported β€” joblib needs it to deserialise the .pkl model files.


πŸ—οΈ System Architecture

Farmer Inputs (soil + weather + location + farm details)
        β”‚
        β”œβ”€β”€β–Ί Model 1 (Random Forest)
        β”‚    model1_npk.pkl + model1_label_encoder.pkl
        β”‚    β†’ probability score per crop (22 crops)
        β”‚
        β”œβ”€β”€β–Ί Model 2 (Regional Scoring Table)
        β”‚    model2_full_scored.csv
        β”‚    β†’ suitability score per crop (33 crops, filtered by district + season)
        β”‚
        └──► Combined Score (60% Model 1 + 40% Model 2)
                    β”‚
                    β–Ό
             Ranked crop list
                    β”‚
                    β”œβ”€β”€β–Ί Budget filter (skip if < 0.1 ha affordable)
                    β”œβ”€β”€β–Ί Risk scoring (weather + market + budget + water)
                    β”œβ”€β”€β–Ί Confidence labelling (based on Model 1 score)
                    └──► "Why" explanation (dynamically generated from real data)
                                β”‚
                                β–Ό
                       HTML output cards (mobile-first UI)

🌱 Crops by Model Coverage

Model 1 β€” Soil/Climate (22 crops): Apple, Banana, Blackgram, Chickpea, Coconut, Coffee, Cotton, Grapes, Jute, Kidneybeans, Lentil, Maize, Mango, Mothbeans, Mungbean, Muskmelon, Orange, Papaya, Pigeonpeas, Pomegranate, Rice, Watermelon

Model 2 β€” Regional (33 crops): Arhar/Tur, Bajra, Banana, Castor Seed, Cotton, Gram, Grapes, Groundnut, Jowar, Linseed, Maize, Mango, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Rice, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds

Crops in both models get a fully blended score. Crops in only one model get a 0 from the missing model but are still shown if the other model scores them highly.


πŸ“Š Data Coverage

Details
State Maharashtra, India
Districts 35
Seasons Kharif (Jun–Sep), Rabi (Oct–Mar), Summer (Apr–Jun), Whole Year
Historical data range 1997 – 2022 (26 years)
Regional data source Maharashtra district-level crop production statistics
Soil model training data 2,200 samples across 22 crops (100 per crop)
Soil model accuracy 99.6% cross-validation (5-fold stratified)

🏷️ Suitability & Risk Labels

Regional Suitability (Model 2 raw score)

Label Score Range Meaning
🟒 Highly Suitable β‰₯ 75% Strong regional farming history
🟑 Moderately Suitable 50–74% Good candidate, moderate history
🟠 Low Suitability 30–49% Possible but limited evidence
πŸ”΄ Not Recommended < 30% Little or no regional history

Overall Risk (combined score)

Label Score Range Action
🟒 Low < 35 Safe choice for this season
🟑 Medium 35–64 Proceed with planning for main risk factor
πŸ”΄ High β‰₯ 65 Consider reducing land allocation

πŸ‘¨β€πŸ’» Built For

Maharashtra farmers and agricultural advisors. Designed to be used on mobile phones in the field β€” large text, simple language, clear YES/NO verdicts, with detailed breakdown available on demand.

Hackathon criteria addressed:

  • βœ… Integration of multiple data inputs (soil, weather, water, location, budget)
  • βœ… Crop suitability prediction via dual-model AI pipeline
  • βœ… Risk-aware scoring with 4-dimensional uncertainty handling
  • βœ… Clear, simple recommendation outputs with plain-language explanations
  • βœ… Farmer-friendly, mobile-responsive interface
  • βœ… Scalable architecture (local model files, no external API dependencies)