crop-adv / README.md
rishiphadale's picture
Update README.md
89afa7b verified
|
Raw
History Blame Contribute Delete
12.3 kB
---
title: Crop Adv
emoji: 🌍
colorFrom: blue
colorTo: yellow
sdk: gradio
sdk_version: 6.8.0
app_file: app.py
pinned: false
license: mit
---
# 🌾 Maharashtra Crop Recommendation System
An AI-powered crop advisory tool that recommends the most suitable crops for a Maharashtra farmer based on their **soil chemistry**, **local weather**, **district & season**, **water supply**, and **budget** β€” and explains *why* each crop is recommended in plain language.
---
## 🧠 How It Works: The Full Pipeline
### Step 1 β€” Two Models Run in Parallel
When you click **Find Best Crops For Me**, two independent models are run simultaneously:
#### Model 1 β€” Soil & Climate AI (Random Forest)
- Takes 7 numerical inputs: N, P, K, Temperature, Humidity, pH, Rainfall
- Runs them through a pre-trained **Random Forest Classifier** (`model1_npk.pkl`)
- Outputs a **probability score (0–1)** for every crop it knows (22 crops)
- Example: `{"rice": 0.42, "maize": 0.28, "cotton": 0.08, ...}`
- The label encoder (`model1_label_encoder.pkl`) maps numeric class indices β†’ crop names
- **Training:** 2,200 samples Γ— 22 crops, 99.6% cross-validation accuracy (5-fold stratified)
#### Model 2 β€” Regional History Scoring (Weighted Formula)
- Looks up the selected **District + Season** combination in `model2_full_scored.csv`
- Returns a pre-computed **suitability score** for each crop based on 26 years (1997–2022) of Maharashtra district-level crop production data
- **No ML here** β€” this is a transparent, rule-based scoring formula:
```
Suitability Score =
0.35 Γ— Crop Frequency % (how often this crop is grown here)
+ 0.25 Γ— Historically Grown Flag (grown in β‰₯50% of years = 1, else 0)
+ 0.20 Γ— Normalised Avg Yield (yield relative to best crop in district)
+ 0.10 Γ— Dominance Rate (share of total district area)
+ 0.05 Γ— Recent Years Grown (grown post-2015 = recency bonus)
+ 0.05 Γ— Yield Stability (lower variance = higher score)
```
- Covers **35 districts Γ— 4 seasons Γ— 33 crops = 1,603 scored combinations**
---
### Step 2 β€” Scores Are Combined
```
Final Score = 0.60 Γ— Model1_probability + 0.40 Γ— Model2_suitability_score
```
- Model 2 scores are **normalised** to 0–1 before combining (divide by max score in that district/season)
- Crops are then **ranked highest β†’ lowest** by Final Score
- Only crops present in `CROP_META` (hardcoded dictionary, see below) are kept
---
### Step 3 β€” Budget & Land Filtering
For each ranked crop:
| Condition | Result |
|---|---|
| `input_cost Γ— land_ha ≀ budget` | `budget_status = "within"` β€” shown fully |
| `input_cost Γ— land_ha > budget` | `affordable_ha = budget / input_cost` β€” shown with partial land warning |
| `affordable_ha < 0.1 ha` | Crop is **skipped entirely** |
---
### Step 4 β€” Risk Scoring (4-Dimensional)
Every crop gets a risk score from 0–100 across four dimensions, then averaged equally:
| Dimension | Weight | How It's Calculated |
|---|---|---|
| Weather Risk | 25% | Penalises extreme temp (<8Β°C or >42Β°C), extreme humidity, and rainfall mismatches vs crop water need |
| Market Risk | 25% | Based on hardcoded demand: High=20, Medium=40, Low=65 |
| Budget Risk | 25% | Ratio of `total_cost / budget_max` mapped to risk bands |
| Water Risk | 25% | Gap between crop's irrigation need vs farmer's available supply |
**Irrigation need levels** (hardcoded per crop):
- `"high"` β€” Rice, Sugarcane, Banana, Grapes, Jute, Apple
- `"medium"` β€” Cotton, Maize, Wheat, Onion, Tomato, etc.
- `"low"` β€” Chickpea, Bajra, Jowar, Pulses, Oilseeds, etc.
**Irrigation supply levels** (farmer input):
- Assured (Canal / River) = 3
- Borewell / Pump = 2
- Seasonal / Rain-fed = 1
- No Irrigation (Dryland) = 0
**Risk label thresholds:**
- Score < 35 β†’ 🟒 **Low**
- Score 35–64 β†’ 🟑 **Medium**
- Score β‰₯ 65 β†’ πŸ”΄ **High**
---
### Step 5 β€” Confidence Labelling
Each result is tagged with an AI confidence level based on **Model 1's raw probability**:
| Model 1 Score | Label | Meaning |
|---|---|---|
| β‰₯ 35% | High confidence | Soil & climate strongly match |
| 12–34% | Moderate confidence | Partial soil match, regional history confirms |
| < 12% | Based on regional history | Driven primarily by Model 2 |
---
### Step 6 β€” "Why We Recommend This" Explanation
Every card generates **5 plain-language bullet points** dynamically from real data β€” nothing is templated or hardcoded:
| Bullet | Data Source |
|---|---|
| Soil fit | `m1_score` from Model 1 probability |
| Season fit | `season_match` boolean (best_season == selected season) |
| Water fit | Irrigation gap calculation (need βˆ’ available) |
| Risk summary | `risk` label + highest `risk_breakdown` dimension |
---
## πŸ“₯ Inputs
### Location & Season
| Input | Type | Options |
|---|---|---|
| District | Dropdown | 35 Maharashtra districts |
| Season | Dropdown | Kharif (Jun–Sep), Rabi (Oct–Mar), Summer (Apr–Jun), Whole Year |
### Farm Details
| Input | Type | Range |
|---|---|---|
| Land (Hectares) | Number | 0.1 – 100 |
| Total Budget (β‚Ή) | Number | β‚Ή1,000+ |
| Water source | Radio | Assured / Borewell / Seasonal / Dryland |
### Soil Test Results (from Krishi Vigyan Kendra or soil lab)
| Input | Range | Unit |
|---|---|---|
| Nitrogen (N) | 0 – 300 | mg/kg |
| Phosphorus (P) | 0 – 300 | mg/kg |
| Potassium (K) | 0 – 300 | mg/kg |
| Soil pH | 3.5 – 9.5 | β€” |
### Local Weather
| Input | Range | Unit |
|---|---|---|
| Avg Temperature | 5 – 50 | Β°C |
| Avg Humidity | 10 – 100 | % |
| Annual Rainfall | 20 – 3000 | mm/year |
---
## πŸ“€ Output: What Each Card Shows
| Element | Source | Notes |
|---|---|---|
| Crop name + medal | Combined score ranking | πŸ₯‡πŸ₯ˆπŸ₯‰ for top 3 |
| Best season | `CROP_META` hardcoded | e.g. "Kharif", "Rabi", "Whole Year" |
| Days to harvest | `CROP_META` hardcoded | Approximate growing period |
| Budget badge | Budget filter logic | βœ… Fits / ⚠️ Partial |
| Risk badge | Risk score calculation | Low / Medium / High |
| "Why we recommend this" | Live data from both models | 4 dynamic bullet points |
| Irrigation verdict | Irrigation gap formula | βœ… / ⚠️ / ❌ |
| Cost overview | `CROP_META` Γ— land_ha | Input cost per ha and total |
| Market demand | `CROP_META` hardcoded | High / Medium / Low |
| Risk breakdown bars | 4-dimensional risk formula | 0–100 per dimension |
| Model confidence | Model 1 probability | High / Moderate / Regional |
| Score breakdown | Both model scores | Soil %, Regional %, Combined % |
---
## πŸ—„οΈ What Is Hardcoded (`CROP_META`)
The following per-crop values are **static estimates** based on Maharashtra Agriculture Department averages. They are not fetched from any live source:
| Field | Description | Example (Rice) |
|---|---|---|
| `input_cost` | Estimated cost to grow 1 hectare (β‚Ή) | β‚Ή30,000 |
| `yield_qtl` | Expected yield in quintals per hectare | 25 qtl/ha |
| `demand` | Market demand category | "High" |
| `days` | Days from sowing to harvest | 120 days |
| `best_season` | Ideal growing season | "Kharif" |
**Full crop list in CROP_META (50 crops):**
Rice, Maize, Chickpea, Kidneybeans, Pigeonpeas, Mothbeans, Mungbean, Blackgram, Lentil, Pomegranate, Banana, Mango, Grapes, Watermelon, Muskmelon, Apple, Orange, Papaya, Coconut, Cotton, Jute, Coffee, Arhar/Tur, Bajra, Castor Seed, Gram, Groundnut, Jowar, Linseed, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds, Cotton(Lint)
> ⚠️ **Note:** `input_cost`, `yield_qtl`, and `demand` are estimates. Actual values vary by year, variety, and farming practices.
---
## πŸ—‚οΈ Files in This Space
| File | Description |
|---|---|
| `app.py` | Main Gradio application β€” all logic lives here |
| `model1_npk.pkl` | Trained Random Forest classifier (soil & climate, 22 crops) |
| `model1_label_encoder.pkl` | Label encoder β€” maps class indices β†’ crop names |
| `model2_full_scored.csv` | Pre-computed regional suitability scores (35 districts Γ— 4 seasons Γ— 33 crops) |
| `requirements.txt` | Python dependencies |
| `README.md` | This file |
---
## πŸ“¦ Requirements
```
gradio>=4.0.0
joblib
numpy
pandas
scikit-learn
```
> ⚠️ `scikit-learn` must be listed even though it is never directly imported β€” `joblib` needs it to deserialise the `.pkl` model files.
---
## πŸ—οΈ System Architecture
```
Farmer Inputs (soil + weather + location + farm details)
β”‚
β”œβ”€β”€β–Ί Model 1 (Random Forest)
β”‚ model1_npk.pkl + model1_label_encoder.pkl
β”‚ β†’ probability score per crop (22 crops)
β”‚
β”œβ”€β”€β–Ί Model 2 (Regional Scoring Table)
β”‚ model2_full_scored.csv
β”‚ β†’ suitability score per crop (33 crops, filtered by district + season)
β”‚
└──► Combined Score (60% Model 1 + 40% Model 2)
β”‚
β–Ό
Ranked crop list
β”‚
β”œβ”€β”€β–Ί Budget filter (skip if < 0.1 ha affordable)
β”œβ”€β”€β–Ί Risk scoring (weather + market + budget + water)
β”œβ”€β”€β–Ί Confidence labelling (based on Model 1 score)
└──► "Why" explanation (dynamically generated from real data)
β”‚
β–Ό
HTML output cards (mobile-first UI)
```
---
## 🌱 Crops by Model Coverage
**Model 1 β€” Soil/Climate (22 crops):**
Apple, Banana, Blackgram, Chickpea, Coconut, Coffee, Cotton, Grapes, Jute, Kidneybeans, Lentil, Maize, Mango, Mothbeans, Mungbean, Muskmelon, Orange, Papaya, Pigeonpeas, Pomegranate, Rice, Watermelon
**Model 2 β€” Regional (33 crops):**
Arhar/Tur, Bajra, Banana, Castor Seed, Cotton, Gram, Grapes, Groundnut, Jowar, Linseed, Maize, Mango, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Rice, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds
**Crops in both models** get a fully blended score. Crops in only one model get a 0 from the missing model but are still shown if the other model scores them highly.
---
## πŸ“Š Data Coverage
| | Details |
|---|---|
| State | Maharashtra, India |
| Districts | 35 |
| Seasons | Kharif (Jun–Sep), Rabi (Oct–Mar), Summer (Apr–Jun), Whole Year |
| Historical data range | 1997 – 2022 (26 years) |
| Regional data source | Maharashtra district-level crop production statistics |
| Soil model training data | 2,200 samples across 22 crops (100 per crop) |
| Soil model accuracy | 99.6% cross-validation (5-fold stratified) |
---
## 🏷️ Suitability & Risk Labels
### Regional Suitability (Model 2 raw score)
| Label | Score Range | Meaning |
|---|---|---|
| 🟒 Highly Suitable | β‰₯ 75% | Strong regional farming history |
| 🟑 Moderately Suitable | 50–74% | Good candidate, moderate history |
| 🟠 Low Suitability | 30–49% | Possible but limited evidence |
| πŸ”΄ Not Recommended | < 30% | Little or no regional history |
### Overall Risk (combined score)
| Label | Score Range | Action |
|---|---|---|
| 🟒 Low | < 35 | Safe choice for this season |
| 🟑 Medium | 35–64 | Proceed with planning for main risk factor |
| πŸ”΄ High | β‰₯ 65 | Consider reducing land allocation |
---
## πŸ‘¨β€πŸ’» Built For
Maharashtra farmers and agricultural advisors. Designed to be used on mobile phones in the field β€” large text, simple language, clear YES/NO verdicts, with detailed breakdown available on demand.
**Hackathon criteria addressed:**
- βœ… Integration of multiple data inputs (soil, weather, water, location, budget)
- βœ… Crop suitability prediction via dual-model AI pipeline
- βœ… Risk-aware scoring with 4-dimensional uncertainty handling
- βœ… Clear, simple recommendation outputs with plain-language explanations
- βœ… Farmer-friendly, mobile-responsive interface
- βœ… Scalable architecture (local model files, no external API dependencies)