The combination of satellite imagery, climate data, and artificial intelligence has been the approach taken by some researchers to create models capable of predicting crop productivity. This type of development has become increasingly necessary in a context of climate change and extreme events, which threaten agricultural productivity and, without adequate forecasting and mitigation measures, can cause losses of billions of dollars.

Photo: Disclosure
It was with this in mind that a group of researchers developed a model capable of estimating the productivity of soybean crops in the Midwest even before harvest. Entitled “Soybean yield estimation in the Brazilian Midwest using Sentinel-2 imageryThe study analyzed data from municipalities in Goiás, Mato Grosso, and Mato Grosso do Sul between the 2019/2020 and 2021/2022 harvests. Using this information and combining it with Sentinel-2 satellite images and machine learning algorithms, the researchers developed a model that achieved 72% accuracy and an average error of less than 302 kg per hectare in estimating soybean productivity.
Soybeans are one of the world's most important agricultural commodities, and Brazil has held the position of the world's largest producer of the grain for at least a decade, according to data from the United States Department of Agriculture (USDA).
In this scenario, soybean cultivation plays a strategic role in the Brazilian economy, both in the domestic market and in foreign trade. In 2024, national production was estimated at 147.38 million tons, cultivated in an area of 46.03 million hectares, of which 46% are concentrated in the Central-West region, according to a survey by the National Supply Company (Conab). Furthermore, data from the Ministry of Agriculture and Livestock (Mapa) indicate that, between 2024 and 2025, Brazilian exports of soybeans, soybean meal, and soybean oil totaled US$$ 52 billion.
Despite its centrality to the Brazilian economy, increasing climate instability has made it increasingly important.

Photo: Caio Inácio
Predicting crop productivity is becoming more difficult. One example is the 2023/2024 harvest, marked by periods of extreme heat and irregular rainfall across much of the country. Initially, Conab projected a soybean harvest of 162 million tons; however, the result was 147.7 million tons, a reduction of approximately 101% compared to initial expectations, according to the report. Extreme heat and agriculture, developed by the Food and Agriculture Organization of the United Nations (FAO). Considering the average soybean prices during the period, this shortfall represented an estimated loss of more than R$ 28 billion for the production sector.
Combination of terrestrial and satellite data
This article is the result of the master's research of Ester de Carvalho Pereira, from the Luiz de Queiroz Higher School of Agriculture of USP (Esalq/USP), which was supervised by professor and researcher Ana Cláudia dos Santos Luciano, from the same institution, and was part of the project “PreCISIA – Crop Prediction by Satellite Image and

Photo: RRRufino
"Artificial Intelligence," funded by the Human Resources Training Program in Strategic Areas (RHAE) of CNPq and coordinated by the company Espectro Ltda. The work was done in collaboration with Michel Eustáquio Dantas Chaves, professor and researcher at the Faculty of Sciences and Engineering of Unesp, Tupã campus, and also integrated researchers from the State University of Londrina and Peking University, in China.
The research combined high-resolution satellite imagery, climate variables, and historical municipal productivity data from IBGE (Brazilian Institute of Geography and Statistics). “The current moment in agriculture is a golden age compared to the past, when data for analysis was lacking. Satellite data allows us to monitor harvests and production cycles, something that until recently was unfeasible, especially at the crop level,” explains Chaves.

At the same time, the researcher points out that the enormous volume of available information poses challenges. "We have a large volume of data ready for analysis, but at the same time, there is the great work of processing and storing it," says the professor.
According to the researchers, in this context it is essential to learn how to work assertively with this large [context].

Photo: Disclosure
The sheer volume of data makes it important to identify which data should be used and over what period of time.
Artificial intelligence played a key role in identifying which variables truly mattered. According to Ester, the study used internal tools of the algorithm itself to measure the weight of each variable in the predictions. "We processed all the variables, ran the model, and one of the questions was to verify which variables had the greatest impact on productivity prediction," she explains.
The results showed that accumulated precipitation, solar radiation, and water deficit were the most important climatic variables for predicting productivity. Among the indicators obtained from satellite images, bands related to infrared and the so-called red edge stood out, a spectral range highly sensitive to the photosynthetic activity of plants.

Photo: Disclosure
Based on this information, the researchers developed six distinct models, each representing a phase of crop development, from 30 to 180 days after planting. At this stage, the objective was to identify which temporal variables exerted the greatest influence on the productivity estimate and, thus, define the most suitable set for constructing the model with the best performance.
In this case, Luciano explains that variables do not refer to information from different sources: all models considered the same sources and types of data, defined in the previous step, but rather to the amount of information in relation to time. "We have the same types of variables in each model, but those that analyzed 150 or 180 days had a larger number of data points because they considered a longer analysis period," he says.
"Considering an initial model that used all the information from the six months studied, we had..."

Photo: Jose Fernando Ogura
"There are approximately 400 variables. But when we create a 30-day model, which has already generated a very interesting result, we work with about 80 variables," adds Ester.
Thus, to ensure the model's efficiency and avoid excessive workload, the group determined the maximum time needed to generate a model with good accuracy. The best performance occurred with the 150-day model, precisely during the soybean grain-filling period, a crucial phase for determining final productivity.
Search for data
Luciano already had prior experience in developing forecasting models for sugarcane cultivation. His work with soybeans arose from a partnership with Espectro, a company specializing in hardware and software development. "Initially, the project was to estimate the productivity of irrigated grain crops, such as soybeans, beans, and corn. Only later did we end up focusing solely on soybeans," he recalls.

Photo: Shutterstock
Prior experience with sugarcane allowed the group to start from an already consolidated base. However, while sugarcane cultivation has a broad and structured database, favoring greater accuracy and ease in model development, the reality of soybean farming is quite different.
Because of this, the work also highlights structural limitations faced by Brazilian science. One of the main difficulties was the lack of detailed data at the rural property level. To overcome this problem, the researchers had to use maps from IBGE (Brazilian Institute of Geography and Statistics) and the MapBiomas project to identify the areas cultivated with soybeans. “The biggest difficulty is having data for that specific location, but we work with what we have,” stated Luciano. “If you want a level of detail so that the producer can actually see their area, it won't meet the needs, but at the government, municipal, and public policy levels, the work applies very well,” he emphasizes.
According to the researchers, the study was not intended to replace field surveys conducted by government agencies.

Photo: Shutterstock
official data is needed, but these analyses should be complemented by continuous large-scale monitoring. Furthermore, the researchers emphasize the importance of integrating different Brazilian public databases, which, according to Chaves, represents one of the sector's main bottlenecks. "If we don't know how to cross-reference the data provided by IBGE with data from Conab, INPE, or Embrapa, among others, we will be losing out as a production chain and as a nation-state," he stresses.
Lack of confidence in national data
According to Luciano, one of the problems that leads to the difficulty of integrating this data is a certain distrust in society regarding data produced by Brazilian institutions, which, for the researcher, represents a devaluation of national science. "We have so many monitoring programs done by us, with Brazilian technology, by Brazilian researchers, that often end up being undervalued here, but are used by other countries," she points out.
40Photo: Disclosure/OPRAs an example, researchers highlight China, the United States, and the European Union, which maintain technologies for monitoring Brazilian crops using Brazilian databases.
Adding to this problem, Chaves points out that many Brazilian monitoring programs suffer from a lack of funding. "The data is useful, accessible, increasingly accurate, and there are more and more people who understand geoprocessing and computing, but there is a lack of adequate investment to expand the return to society," he mentions.
Funding limitations have led to the closure of monitoring programs such as Cafésat and Canasat, which were maintained by the Brazilian Space Agency (INPE), and threaten the continuity of others such as... TerraClass and the Prodes.
