Applied AI for plant breeding
On 17 February 2026, Dr Maryna Kuzmenko — founder of Petiole and Petiole Pro — spoke at the 2026 Soybean Breeders' Workshop, hosted by Brian Little and his team at the Center for Applied Genetic Technologies, University of Georgia. The talk was titled Applied AI for Soybean Breeders: Tools, Workflows, Use Cases, and it was delivered virtually from a rainy Peterborough, UK.
The core argument: most published lists of "AI for soybean" describe what AI could do. Applied AI is the much shorter list of what actually works in a breeding programme today — and nearly all of it is visual, smartphone-based, and stage-specific. Counting germinated seeds, measuring leaf area, quantifying leaf damage, counting flowers and pods, counting root nodules, and running post-harvest seed QC are real. Everything else is still a research direction.
This long-read reconstructs the full session: what separates AI from applied AI, which imaging tool belongs at each soybean growth stage, why drones cannot count soybean flowers, where satellite imagery stops being useful, what the Kansas State nodule-counting collaboration produced, and why the number on its own is never the deliverable.
- Ask a chatbot how to use AI in soybean research and you get five to ten directions. Try them and one, maybe two, actually work today. That gap is what "applied AI" means.
- The tool follows the growth stage, not the other way round: RGB smartphone for germination, leaf area and seed QC; drones for emergence, lodging and maturity; thermal for water stress; satellite for almost nothing a breeder needs.
- Drones cannot count soybean flowers or pods — they are hidden under the canopy, and white or pale flowers are hard for a computer vision model to separate from leaves even with good labelling.
- An existing model is not a working model. The germination module trained on tree seedlings failed on soybean trays. Relevant training data is the bottleneck, not the algorithm.
- A count with no context evaporates. Variety, sowing date, fertiliser, irrigation, soil test — the number only becomes data when it travels with the journey of growing.
The workshop, and why this talk happened
The 2026 Soybean Breeders' Workshop ran as a hybrid event out of the University of Georgia. The invitation arrived by an unusually honest route: Brian Little had been researching handheld devices that could scan and analyse things in the field, went down a YouTube rabbit hole, and found Petiole Pro.
Petiole is a UK-based company, nine years old, whose flagship free mobile application turns smartphone and drone imagery into AI-powered crop insights. The platform has been cited in more than 110 research papers, and the work has been supported by three grants from the UK and EU governments. Three of those papers reference Petiole Pro directly as the data-collection tool inside soybean experiments.



AI vs. applied AI for soybean
Ask any search engine or chatbot "what is AI for soybeans?" and you will get an answer immediately. The answer will be fluent, structured, and mostly aspirational — a blend of summarisation, ideation, and a little dreaming. It will promise accelerated breeding, real-time pest detection, yield prediction with over 90% accuracy, and irrigation optimisation.

The problem is not accuracy. The problem is applicability. Ask the follow-up question — "how can I use existing tools for my soybean research?" — and you still get five or ten directions. Try to use them, and you end up with one, maybe two, genuinely working cases.
So this talk is not about general, all-purpose, all-directional AI for soybean breeding. It is about what is practical and applicable. Said precisely, the topic is applied, visual, smartphone-based AI for soybeans — because that is what nine years of work suggests holds the biggest promise for practical application.
That framing matters for a specific reason. Researchers already know what is possible in front of very expensive equipment and very expensive cameras. The more useful question — for research work and for extension work alike — is what you can do with an ordinary smartphone in your hand.

The technology stack behind every tool
Every tool in the talk sits on the same core stack, and it is worth naming the parts because the words get used loosely:
- Computer vision — the technological eyes. How a computer sees an image and understands what is on it.
- Machine learning / deep learning — deep learning is a part of machine learning. This is what takes the data from the imagery, processes it, and finds the patterns inside what has been collected.
- Remote sensing (optional) — indices, most famously NDVI, used to understand deficiency in crops.
- Sensors (optional) — a nice-to-have that gives you extra layers of data to support a claim: environmental monitoring, and how conditions map onto a specific stage.

Stage 1 — Germination
The job to be done at germination is to assess percentage germination, speed and vigour, and uniformity. In practice that breaks into two tasks: count the germinated seeds, measure radicle length, classify abnormal seedlings — and then build the germination curve automatically.
The tool is an ordinary RGB camera. A smartphone camera is an RGB camera, so nothing extra is needed. From the count you get germinated and non-germinated cells, and the germination rate falls out automatically.
The bonus layer is environmental: temperature and moisture sensors let you link germination outcomes to the conditions that produced them — which matters when trays sit in different places and your experiment conditions genuinely diversify.

Here is what automation looks like in practice. The imagery below is tree seedlings, not soybeans — but the process runs identically on any tray-based crop, and it is shown to demonstrate what becomes possible for soybean germination trays. The beauty of it is not the algorithm. It is that nobody stands in front of the trays counting with their eyes and their fingers and writing into a notepad.


And here is where the honesty comes in. One Petiole Pro user sent us their own soybean germination tray and asked whether our module could be applied to it. The answer was: you can try — but look at the two images side by side.

The computer vision algorithm sees that difference too — and it is lost. It does not understand what we want from it. The model already knows the concept: there are cells, there may be seedlings inside, no seedling means not germinated. The challenge is not the algorithm. It is providing relevant data to train and adjust it. As with every AI model, bring the relevant data, train the model, and performance and accuracy can be improved and boosted.
That distinction — an existing model versus a working model — is the single most useful thing a breeder can take from this section. If you want to go deeper on the mechanics of automated germination counting, we wrote it up separately in Beyond the Sprout: AI germination counting.
Stage 2 — Sprout and emergence
The job to be done shifts to emergence percentage, stand count, gaps, and early stress. Two approaches: emergence counting at plot scale, and early vigour trends via indices.
You can use a smartphone. You can use a drone — RGB, or multispectral if you have it. But the most interesting method came from a researcher running a low-cost experiment: he attached a smartphone to a selfie stick, pointed it down at the plants, and walked along the row recording video. It is not innovation in any grand sense — nothing about it is new — and that is exactly why it works. It scaled, because walking scales, and the video processed cleanly into counts.

At this stage you can also link the germination and count data to weather-condition data and build the dependency. And you can start picking up broader trends — early disease signal, or the presence of chlorosis. But what we discovered is that the scope widens dramatically once leaf growth is present and there are at least a few leaves to work with. That is where it gets interesting.
Stage 3 — Leaf growth and the vegetative stage
Now the job to be done is canopy cover, vigour, LAI proxies, nutrient stress, weeds, and pest or disease onset. Two workhorses — remote sensing with drone multispectral plus drone RGB, and in-field sampling with a mobile application. Satellite imagery sits in the nice-to-have column, with three caveats attached: price, time, and accuracy.

RGB, multispectral, and the purple-light trap
Multispectral imaging can detect disease earlier than RGB because it captures wavelengths beyond visible light. It is worth seeing what that actually looks like, because there is a trap here that catches people out.

Multispectral purple and vertical-farming LED purple are not the same. If you capture an image under LED grow lights, you are still processing it as ordinary RGB imagery — and you cannot get early disease detection in those conditions. Multispectral can. (Whether multispectral capture under LED is worth testing is an open and rather interesting question.)
Leaf area measurement
Once there are leaves, the first genuinely applicable smartphone task is leaf area measurement. This is not new science — it is a standard procedure. What changes is that a manual task becomes a digital one, which brings accuracy and, more importantly, scalability.

The scalability point is concrete. Manual leaf area assessment with ImageJ or grid paper can consume a day. With the mobile app, some researchers have taken thousands of photos, because they have a large selection of plants and need to measure all of them — and in that case more data directly means better accuracy of results. The method is simply: photograph the leaf in front of a calibration plate, or place the leaf on it.

The calibration plates come in several sizes, driven by requests from growers. The smallest was developed specifically for a researcher in Colombia measuring leaf area in natural conditions high in the mountains — the app works offline, without internet, which is what made that possible at all. If you want the detail on why the plate is there and what happens without it, we covered that in Do I need a calibration plate for Petiole Pro?, and the capture technique in how to photograph leaves for accurate leaf area measurement in the field.

There are thousands, maybe millions, of scenarios where leaf area is a useful parameter: any agricultural input trial, any attempt to assess the impact of an environmental change on a specific variety. This is one method of automating it.
Early-stage change detection: the cucumber experiment
A farmer with six rows of plants applied a different treatment to each row. Visually he could tell that some rows were greener than others. His question was the right one: how can I prove it?
We set a greenness threshold and built a model around it. He shot video; the pipeline segmented each leaf out of the frames; each leaf was then compared against the threshold. He had never had numbers before. We returned six pages of metrics — per-leaf, per-row, against overall, against average — because once the data is good, extraction is not the hard part.

Quantifying disease and pest damage
Leaf area is straightforward: count it. Disease is harder, because the question is not "is it there" but "how much, and which way is it moving?"
One of our UK government grants funded exactly this: an AI system for early-stage detection and monitoring of apple scab and apple canker for UK farmers. Two tasks — detect the lesions, which is difficult early, and then measure them. From the measurement you get a threshold, and from repeated measurement you get the dynamic: is the disease declining or not?

A deliberate design choice: report percentages, not raw square centimetres. Farmers do not have time to work out whether 24 cm² is big or small. They love percentage — so deliver the number that is actually useful for the decision.
The same measurement method covers pest damage. There are two distinct kinds of damaged area, and they need different treatment: coloured or discoloured spots on the leaf, and physically missing, chewed-up area.

For the chewed case, you can quantify the area that remains after the pest has eaten — and, more usefully, calculate which area has been eaten and report it as a percentage. That is the number that tells you about impact.

We work with soft fruit growers counting pests on sticky yellow traps, but for soybean the more relevant experience is the damaged-area measurement above. The soybean case study is in the pipeline. If you want the full method behind damage quantification — segmentation, colour indices, and why percentages and absolute cm² are both worth having — it is written up in AI assessment of necrotic leaf area damage.
Stage 4 — Flowering, and why drones fail here
Flowering is where yield lives, so the job to be done is flowering date, uniformity, flower density, and stress impacts on set. In-field sampling with a mobile application comes first; drone RGB second; thermal imagery third. Phenology models — predicting the flowering window from weather plus planting date — are a nice-to-have for trial planning.

Here is the part that is worth saying plainly, because it cuts against the direction of travel in a lot of phenotyping conversations. Drones cannot properly capture soybean flowers. Two reasons, and they compound:
- The flowers are hidden. They sit below the leaf coverage. A camera looking down sees leaves.
- Colour separation is hard. If the soybean has purple flowers, you at least have a difference to work with. If they are white or yellowish, it is a genuinely difficult problem for a computer vision algorithm even with good labelling — labelling being the step where you tell the model, on the photo, this is a flower and this is a leaf.
Add drone motion blur to that. To reduce it you need slower speed and lower altitude — but fly lower and you still cannot see the flowers under the top leaves. It is not commercially viable and it is not practically viable to use an unmanned aerial vehicle for counting soybean flowers.
What works instead: handheld devices collecting imagery at plant level, not above it. Or a smartphone attached to a quad bike or tractor, collecting at canopy height. Data from the level of the plants is more insightful than data from the top. And on a huge field, sample — then scope the count with statistics.
Get the raw flower count first with computer vision. Then add phenology models on top for a richer outcome. But the count is the foundation; everything else is layered onto it.
Stage 5 — Pod fill and maturity
Pod fill inherits the flowering problem exactly: the pods are the same colour as the foliage, and they sit below the leaves. The job to be done is pod set success, stress during fill, biomass, and yield proxies — and once again in-field phenotyping outperforms remote-sensing phenotyping.

Thermal imagery earns its place at this stage and earlier. From the first leaves onward, thermal cameras on drones detect water stress and flag areas of interest. Drones cannot count your pods or your flowers — but they can find your stress, and that method transfers to other crops.

Stage 6 — Post-harvest seed QC
Everything so far has been in-field or in-greenhouse. Post-harvest is different, and it is one of the clear success stories — because it happens in a controlled environment. You are not dependent on temperature or weather. You have one lighting setup and one data-collection protocol.
The job to be done: count, size distribution, cracks, discoloration, mould, impurities, varietal purity, and viability proxies. Two routes, which stack rather than compete:
- Mobile application, visual QC. Count, size, presence or absence of defects, visible cracking, and any problems on the sample.
- NIR / hyperspectral / X-ray. For internal state. Computer vision and RGB imagery are good at what we can see; they say nothing about the chemical profile of a seed or its internal problems. Proteins and oils are discoverable this way — but the equipment is expensive.
At production level, this becomes automated QC machines and sorters, again built on computer vision.

Three workflows: satellite, drone, smartphone
Satellite — and why it barely featured
Satellite imagery came up once, briefly, and mostly to explain its absence. Plant breeders are interested in very specific results from specific experiments. Satellite is for policy and for understanding the overall scope of production. The resolution simply does not deliver the insights a breeding programme needs.

Drone RGB — where drones genuinely win
Drones are excellent, and the talk was clear about where: emergence counts, lodging, and maturity. Lodging in particular is a case where handheld simply cannot compete — you need to go up to see the affected areas at the later stage. Maturity, being a colour-difference problem across a field, is also a good fit.

Mobile devices — the through-line
Germination. In-field scouting: plant health, flower and yield count, pest and disease pressure. Post-harvest seed visual QC. And then a bonus category that nobody asks for and everybody needs: operational routine.

That standard-deviation view is worth pausing on. We do not just count and show the count — we show how the seeds are distributed: how many sit in the average range, how many at the smallest, how many at the biggest. It is the difference between a number and a characterisation.

Operational excellence: the part nobody asks for
This was the quiet heart of the talk. Years of work with breeders and growers taught us that getting the number is not the task. You always need to attach something to the number: which variety, when it was sown, when fertiliser or any agricultural input was applied. It is never a story of one number. It is a journey of growing, and the context has to attach smoothly — otherwise the count is done, and then it disappears, because the next stage has already started.

The pain this solves is embarrassingly ordinary. A grower with three experiments and four fertiliser applications works in the field, comes back to the office, and tries to transfer his notes into an Excel table. It is time-consuming. Sometimes the notes are lost. Sometimes they have been out in the rain and are not very good. And the historical record ends up somewhere in Google Drive, to be hunted for later.

There is a second reason to keep it together beyond not losing it. Once all the data is in one place, you can feed it into an AI algorithm that provides more insight — because sometimes we miss things, and summarising data is exactly what these systems are good at.

An unexpected one: nozzle output
One of the soybean papers citing Petiole Pro was nominally about leaf area — but the researchers' real task was assessing nozzle setup on a sprayer. They used water-sensitive paper and needed to quantify the footprint of each nozzle at each setting. That is not a breeding application, but applying agricultural chemicals correctly is an important job, and the same measurement machinery works: quantify the print output instead of eyeballing it, and get at least a percentage.


Nodules, stomata, and models trained on demand
Two final pieces of science, both of which started as somebody's problem rather than our roadmap.
Stomata counting was brought to us by wheat breeders trying to assess their success, for whom manual counting was punishing. It is not a soybean task, but it may be useful — and it now exists as both a web and a mobile application.

Nodule counting is a soybean success story, and it came from a researcher at Kansas State University. He got in touch with a fair complaint: you measure leaves, but I work with roots, and counting these nodules takes my whole day when I am busy with other tasks — can you help? It was an interesting task, so we took it. That was around three years ago, before the AI boom, when computer vision was nothing like as popular an application as it is now.


This is what "AI on demand" means in practice, and it closes the loop back to the germination failure. The germination model did not work on soybean because it had not seen soybean. But that is a solvable problem: if something is already developed, we can go deeper on your imagery and build it together.
Watch the full talk
The complete session — all six stages, the tool comparisons, the failure cases, and the audience questions — is on YouTube.
If the basics underneath all of this are what you actually want — what machine learning is, how deep learning differs from it, what separates traditional farming from smart farming — there is a free introductory course on Udemy. AI moves fast; agriculture moves at the speed of seasons and biological cycles, which is not the same speed at all, and it is worth reminding ourselves of the fundamentals.

Four themes from the wider programme
The talk was one session in a full programme, and what stayed with us was not only our own. Four directions felt especially relevant across the workshop:
- Speed and throughput are no longer a nice-to-have. Winter nurseries, smarter greenhouse workflows, technician-led efficiency — that is how cycle time is measured in reality. Otherwise routine eats the time.
- High-throughput phenotyping was discussed mainly through drone imaging for maturity and field performance. We were pitching handheld tools — but with genuine (and mutual) affection for drones, it is good to see this data-collection method treated as a repeatable signal for better in-field decisions.
- Durable resistance is gaining power. Soybean Cyst Nematode and disease pressure remain relentless, and there was a lot of practical thinking around stacking, validation, and keeping resistance useful over time rather than winning a single season. Everyone is actively scouting novel resistance sources, including wild soybean resources.
- Quality traits are moving earlier in the pipeline. Oil profile, protein, composition analytics as usual — but now scaling fatty-acid profiling (FAME/GC) and other compositional screens earlier, not only at the end.
And the cherry on top — or better, a bunch of beautiful mature soybean pods: digitalisation is slowly but confidently wrapping more and more aspects of soybean breeding. Innovation rarely arrives as a big reveal. It arrives as dozens of small, disciplined, stacked improvements until the whole pipeline moves faster.
If you work in wheat, barley, oilseed rape, pulses, vegetables or berries, the parallel themes hold. The tools differ and the constraints differ, but the directions are shared.
Petiole Pro for soybean breeding
Bring your imagery, and let's see what works
The Petiole Pro mobile app is free on the Play Store and works offline — measure leaf area, count seeds, and assess damage from your phone. If you have a task that needs a model trained on your own soybean imagery — germination trays, nodules, pods, flowers — that is exactly the conversation we want. Write to us at [email protected].
Building a grant application that needs phenotyping, seed QC, or damage measurement? We are a UK-based company with Innovate UK plant-science grants behind us and are glad to join as a consortium partner.

Frequently asked questions
What is the difference between AI and applied AI for soybean breeding?
AI for soybean, as usually described, is a list of possible directions — accelerated breeding, yield prediction, pest detection, irrigation optimisation. Applied AI is the much shorter list of what works in a breeding programme today. Ask a chatbot how to use existing tools for soybean research and you get five to ten directions; try them and one or two are genuinely usable. Applied AI for soybean is overwhelmingly visual, smartphone-based, and tied to a specific growth stage.
Can AI count soybean germination from a photo of a tray?
Yes in principle — an RGB smartphone camera is enough to count germinated and empty cells and compute the germination rate automatically. But a model trained on another crop's trays will fail on soybean: in the workshop demo, a module that scored 240 tree-seedling cells at 92.9% germination could not read a soybean tray, because the plug geometry and seedling architecture differ. The model concept exists; the bottleneck is relevant soybean training data.
Can drones count soybean flowers and pods?
No, not practically. Soybean flowers and pods sit below the leaf canopy, so a downward-looking camera mostly sees leaves. White or yellowish flowers are also hard for a computer vision model to separate from foliage even with good labelling, and drone motion adds blur. Handheld devices at plant level — or a smartphone mounted on a quad bike or tractor — are more insightful than imaging from above.
What are drones actually good for in soybean?
Emergence counts, lodging, and maturity. Lodging in particular is hard to see with handheld devices — you need altitude to spot the affected areas late in the season. Maturity is a colour-difference problem across a field, which suits drone RGB well. Thermal cameras on drones are also effective for water-stress detection from the first leaves onward.
Is satellite imagery useful for soybean breeding?
Rarely. Satellite imagery suits policy work and understanding the overall scope of production. Breeders need specific results from specific experiments, and the resolution does not deliver them: a poor soybean stand visible in drone imagery at 0.75 inch per pixel remains identifiable at 5 m/pixel but with much less detail, and at 10 m/pixel the detail is lost and the utility for decision support diminished.
What can a smartphone measure on soybean that expensive equipment usually does?
Leaf area (with a calibration plate, in real cm², offline), germination counts, greenness and early chlorosis against a threshold, disease lesion severity as a percentage of leaf area, pest-chewed area, flower and pod counts at plant level, root nodule counts, and post-harvest seed count with size distribution and defect screening. What it cannot do is see inside a seed — chemical profile, protein and oil content need NIR, hyperspectral or X-ray.
Why report leaf damage as a percentage rather than in cm²?
Because growers do not have time to work out whether 24 cm² is a lot or a little. Percentage is immediately interpretable and drives the decision. For scientific work, both are worth having: the percentage gives you the relative severity, and calibrated cm² gives you the absolute measurement.
Why does a count need context like variety and fertiliser attached to it?
Because a number alone disappears. Working with breeders and growers showed that a count only becomes data when it travels with the journey of growing — variety, sowing date, agricultural inputs, irrigation, soil tests. Field notes transferred later into Excel get lost, rained on, or scattered across drives. Keeping everything in one place also makes it possible to feed the whole record into an AI system for summarisation and further insight.