Spaces:
Sleeping
Sleeping
| title: Crop Adv | |
| emoji: π | |
| colorFrom: blue | |
| colorTo: yellow | |
| sdk: gradio | |
| sdk_version: 6.8.0 | |
| app_file: app.py | |
| pinned: false | |
| license: mit | |
| # πΎ Maharashtra Crop Recommendation System | |
| An AI-powered crop advisory tool that recommends the most suitable crops for a Maharashtra farmer based on their **soil chemistry**, **local weather**, **district & season**, **water supply**, and **budget** β and explains *why* each crop is recommended in plain language. | |
| --- | |
| ## π§ How It Works: The Full Pipeline | |
| ### Step 1 β Two Models Run in Parallel | |
| When you click **Find Best Crops For Me**, two independent models are run simultaneously: | |
| #### Model 1 β Soil & Climate AI (Random Forest) | |
| - Takes 7 numerical inputs: N, P, K, Temperature, Humidity, pH, Rainfall | |
| - Runs them through a pre-trained **Random Forest Classifier** (`model1_npk.pkl`) | |
| - Outputs a **probability score (0β1)** for every crop it knows (22 crops) | |
| - Example: `{"rice": 0.42, "maize": 0.28, "cotton": 0.08, ...}` | |
| - The label encoder (`model1_label_encoder.pkl`) maps numeric class indices β crop names | |
| - **Training:** 2,200 samples Γ 22 crops, 99.6% cross-validation accuracy (5-fold stratified) | |
| #### Model 2 β Regional History Scoring (Weighted Formula) | |
| - Looks up the selected **District + Season** combination in `model2_full_scored.csv` | |
| - Returns a pre-computed **suitability score** for each crop based on 26 years (1997β2022) of Maharashtra district-level crop production data | |
| - **No ML here** β this is a transparent, rule-based scoring formula: | |
| ``` | |
| Suitability Score = | |
| 0.35 Γ Crop Frequency % (how often this crop is grown here) | |
| + 0.25 Γ Historically Grown Flag (grown in β₯50% of years = 1, else 0) | |
| + 0.20 Γ Normalised Avg Yield (yield relative to best crop in district) | |
| + 0.10 Γ Dominance Rate (share of total district area) | |
| + 0.05 Γ Recent Years Grown (grown post-2015 = recency bonus) | |
| + 0.05 Γ Yield Stability (lower variance = higher score) | |
| ``` | |
| - Covers **35 districts Γ 4 seasons Γ 33 crops = 1,603 scored combinations** | |
| --- | |
| ### Step 2 β Scores Are Combined | |
| ``` | |
| Final Score = 0.60 Γ Model1_probability + 0.40 Γ Model2_suitability_score | |
| ``` | |
| - Model 2 scores are **normalised** to 0β1 before combining (divide by max score in that district/season) | |
| - Crops are then **ranked highest β lowest** by Final Score | |
| - Only crops present in `CROP_META` (hardcoded dictionary, see below) are kept | |
| --- | |
| ### Step 3 β Budget & Land Filtering | |
| For each ranked crop: | |
| | Condition | Result | | |
| |---|---| | |
| | `input_cost Γ land_ha β€ budget` | `budget_status = "within"` β shown fully | | |
| | `input_cost Γ land_ha > budget` | `affordable_ha = budget / input_cost` β shown with partial land warning | | |
| | `affordable_ha < 0.1 ha` | Crop is **skipped entirely** | | |
| --- | |
| ### Step 4 β Risk Scoring (4-Dimensional) | |
| Every crop gets a risk score from 0β100 across four dimensions, then averaged equally: | |
| | Dimension | Weight | How It's Calculated | | |
| |---|---|---| | |
| | Weather Risk | 25% | Penalises extreme temp (<8Β°C or >42Β°C), extreme humidity, and rainfall mismatches vs crop water need | | |
| | Market Risk | 25% | Based on hardcoded demand: High=20, Medium=40, Low=65 | | |
| | Budget Risk | 25% | Ratio of `total_cost / budget_max` mapped to risk bands | | |
| | Water Risk | 25% | Gap between crop's irrigation need vs farmer's available supply | | |
| **Irrigation need levels** (hardcoded per crop): | |
| - `"high"` β Rice, Sugarcane, Banana, Grapes, Jute, Apple | |
| - `"medium"` β Cotton, Maize, Wheat, Onion, Tomato, etc. | |
| - `"low"` β Chickpea, Bajra, Jowar, Pulses, Oilseeds, etc. | |
| **Irrigation supply levels** (farmer input): | |
| - Assured (Canal / River) = 3 | |
| - Borewell / Pump = 2 | |
| - Seasonal / Rain-fed = 1 | |
| - No Irrigation (Dryland) = 0 | |
| **Risk label thresholds:** | |
| - Score < 35 β π’ **Low** | |
| - Score 35β64 β π‘ **Medium** | |
| - Score β₯ 65 β π΄ **High** | |
| --- | |
| ### Step 5 β Confidence Labelling | |
| Each result is tagged with an AI confidence level based on **Model 1's raw probability**: | |
| | Model 1 Score | Label | Meaning | | |
| |---|---|---| | |
| | β₯ 35% | High confidence | Soil & climate strongly match | | |
| | 12β34% | Moderate confidence | Partial soil match, regional history confirms | | |
| | < 12% | Based on regional history | Driven primarily by Model 2 | | |
| --- | |
| ### Step 6 β "Why We Recommend This" Explanation | |
| Every card generates **5 plain-language bullet points** dynamically from real data β nothing is templated or hardcoded: | |
| | Bullet | Data Source | | |
| |---|---| | |
| | Soil fit | `m1_score` from Model 1 probability | | |
| | Season fit | `season_match` boolean (best_season == selected season) | | |
| | Water fit | Irrigation gap calculation (need β available) | | |
| | Risk summary | `risk` label + highest `risk_breakdown` dimension | | |
| --- | |
| ## π₯ Inputs | |
| ### Location & Season | |
| | Input | Type | Options | | |
| |---|---|---| | |
| | District | Dropdown | 35 Maharashtra districts | | |
| | Season | Dropdown | Kharif (JunβSep), Rabi (OctβMar), Summer (AprβJun), Whole Year | | |
| ### Farm Details | |
| | Input | Type | Range | | |
| |---|---|---| | |
| | Land (Hectares) | Number | 0.1 β 100 | | |
| | Total Budget (βΉ) | Number | βΉ1,000+ | | |
| | Water source | Radio | Assured / Borewell / Seasonal / Dryland | | |
| ### Soil Test Results (from Krishi Vigyan Kendra or soil lab) | |
| | Input | Range | Unit | | |
| |---|---|---| | |
| | Nitrogen (N) | 0 β 300 | mg/kg | | |
| | Phosphorus (P) | 0 β 300 | mg/kg | | |
| | Potassium (K) | 0 β 300 | mg/kg | | |
| | Soil pH | 3.5 β 9.5 | β | | |
| ### Local Weather | |
| | Input | Range | Unit | | |
| |---|---|---| | |
| | Avg Temperature | 5 β 50 | Β°C | | |
| | Avg Humidity | 10 β 100 | % | | |
| | Annual Rainfall | 20 β 3000 | mm/year | | |
| --- | |
| ## π€ Output: What Each Card Shows | |
| | Element | Source | Notes | | |
| |---|---|---| | |
| | Crop name + medal | Combined score ranking | π₯π₯π₯ for top 3 | | |
| | Best season | `CROP_META` hardcoded | e.g. "Kharif", "Rabi", "Whole Year" | | |
| | Days to harvest | `CROP_META` hardcoded | Approximate growing period | | |
| | Budget badge | Budget filter logic | β Fits / β οΈ Partial | | |
| | Risk badge | Risk score calculation | Low / Medium / High | | |
| | "Why we recommend this" | Live data from both models | 4 dynamic bullet points | | |
| | Irrigation verdict | Irrigation gap formula | β / β οΈ / β | | |
| | Cost overview | `CROP_META` Γ land_ha | Input cost per ha and total | | |
| | Market demand | `CROP_META` hardcoded | High / Medium / Low | | |
| | Risk breakdown bars | 4-dimensional risk formula | 0β100 per dimension | | |
| | Model confidence | Model 1 probability | High / Moderate / Regional | | |
| | Score breakdown | Both model scores | Soil %, Regional %, Combined % | | |
| --- | |
| ## ποΈ What Is Hardcoded (`CROP_META`) | |
| The following per-crop values are **static estimates** based on Maharashtra Agriculture Department averages. They are not fetched from any live source: | |
| | Field | Description | Example (Rice) | | |
| |---|---|---| | |
| | `input_cost` | Estimated cost to grow 1 hectare (βΉ) | βΉ30,000 | | |
| | `yield_qtl` | Expected yield in quintals per hectare | 25 qtl/ha | | |
| | `demand` | Market demand category | "High" | | |
| | `days` | Days from sowing to harvest | 120 days | | |
| | `best_season` | Ideal growing season | "Kharif" | | |
| **Full crop list in CROP_META (50 crops):** | |
| Rice, Maize, Chickpea, Kidneybeans, Pigeonpeas, Mothbeans, Mungbean, Blackgram, Lentil, Pomegranate, Banana, Mango, Grapes, Watermelon, Muskmelon, Apple, Orange, Papaya, Coconut, Cotton, Jute, Coffee, Arhar/Tur, Bajra, Castor Seed, Gram, Groundnut, Jowar, Linseed, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds, Cotton(Lint) | |
| > β οΈ **Note:** `input_cost`, `yield_qtl`, and `demand` are estimates. Actual values vary by year, variety, and farming practices. | |
| --- | |
| ## ποΈ Files in This Space | |
| | File | Description | | |
| |---|---| | |
| | `app.py` | Main Gradio application β all logic lives here | | |
| | `model1_npk.pkl` | Trained Random Forest classifier (soil & climate, 22 crops) | | |
| | `model1_label_encoder.pkl` | Label encoder β maps class indices β crop names | | |
| | `model2_full_scored.csv` | Pre-computed regional suitability scores (35 districts Γ 4 seasons Γ 33 crops) | | |
| | `requirements.txt` | Python dependencies | | |
| | `README.md` | This file | | |
| --- | |
| ## π¦ Requirements | |
| ``` | |
| gradio>=4.0.0 | |
| joblib | |
| numpy | |
| pandas | |
| scikit-learn | |
| ``` | |
| > β οΈ `scikit-learn` must be listed even though it is never directly imported β `joblib` needs it to deserialise the `.pkl` model files. | |
| --- | |
| ## ποΈ System Architecture | |
| ``` | |
| Farmer Inputs (soil + weather + location + farm details) | |
| β | |
| ββββΊ Model 1 (Random Forest) | |
| β model1_npk.pkl + model1_label_encoder.pkl | |
| β β probability score per crop (22 crops) | |
| β | |
| ββββΊ Model 2 (Regional Scoring Table) | |
| β model2_full_scored.csv | |
| β β suitability score per crop (33 crops, filtered by district + season) | |
| β | |
| ββββΊ Combined Score (60% Model 1 + 40% Model 2) | |
| β | |
| βΌ | |
| Ranked crop list | |
| β | |
| ββββΊ Budget filter (skip if < 0.1 ha affordable) | |
| ββββΊ Risk scoring (weather + market + budget + water) | |
| ββββΊ Confidence labelling (based on Model 1 score) | |
| ββββΊ "Why" explanation (dynamically generated from real data) | |
| β | |
| βΌ | |
| HTML output cards (mobile-first UI) | |
| ``` | |
| --- | |
| ## π± Crops by Model Coverage | |
| **Model 1 β Soil/Climate (22 crops):** | |
| Apple, Banana, Blackgram, Chickpea, Coconut, Coffee, Cotton, Grapes, Jute, Kidneybeans, Lentil, Maize, Mango, Mothbeans, Mungbean, Muskmelon, Orange, Papaya, Pigeonpeas, Pomegranate, Rice, Watermelon | |
| **Model 2 β Regional (33 crops):** | |
| Arhar/Tur, Bajra, Banana, Castor Seed, Cotton, Gram, Grapes, Groundnut, Jowar, Linseed, Maize, Mango, Moong (Green Gram), Niger Seed, Onion, Other Cereals, Other Kharif Pulses, Other Rabi Pulses, Other Summer Pulses, Ragi, Rapeseed & Mustard, Rice, Safflower, Sesamum, Small Millets, Soyabean, Sugarcane, Sunflower, Tobacco, Tomato, Urad, Wheat, Other Oilseeds | |
| **Crops in both models** get a fully blended score. Crops in only one model get a 0 from the missing model but are still shown if the other model scores them highly. | |
| --- | |
| ## π Data Coverage | |
| | | Details | | |
| |---|---| | |
| | State | Maharashtra, India | | |
| | Districts | 35 | | |
| | Seasons | Kharif (JunβSep), Rabi (OctβMar), Summer (AprβJun), Whole Year | | |
| | Historical data range | 1997 β 2022 (26 years) | | |
| | Regional data source | Maharashtra district-level crop production statistics | | |
| | Soil model training data | 2,200 samples across 22 crops (100 per crop) | | |
| | Soil model accuracy | 99.6% cross-validation (5-fold stratified) | | |
| --- | |
| ## π·οΈ Suitability & Risk Labels | |
| ### Regional Suitability (Model 2 raw score) | |
| | Label | Score Range | Meaning | | |
| |---|---|---| | |
| | π’ Highly Suitable | β₯ 75% | Strong regional farming history | | |
| | π‘ Moderately Suitable | 50β74% | Good candidate, moderate history | | |
| | π Low Suitability | 30β49% | Possible but limited evidence | | |
| | π΄ Not Recommended | < 30% | Little or no regional history | | |
| ### Overall Risk (combined score) | |
| | Label | Score Range | Action | | |
| |---|---|---| | |
| | π’ Low | < 35 | Safe choice for this season | | |
| | π‘ Medium | 35β64 | Proceed with planning for main risk factor | | |
| | π΄ High | β₯ 65 | Consider reducing land allocation | | |
| --- | |
| ## π¨βπ» Built For | |
| Maharashtra farmers and agricultural advisors. Designed to be used on mobile phones in the field β large text, simple language, clear YES/NO verdicts, with detailed breakdown available on demand. | |
| **Hackathon criteria addressed:** | |
| - β Integration of multiple data inputs (soil, weather, water, location, budget) | |
| - β Crop suitability prediction via dual-model AI pipeline | |
| - β Risk-aware scoring with 4-dimensional uncertainty handling | |
| - β Clear, simple recommendation outputs with plain-language explanations | |
| - β Farmer-friendly, mobile-responsive interface | |
| - β Scalable architecture (local model files, no external API dependencies) |