Home » Research » Artificial Intelligence and Drones in Monitoring and Maintaining Infrastructure and Urban Road Networks

Artificial Intelligence and Drones in Monitoring and Maintaining Infrastructure and Urban Road Networks

One asphalt road read through layered planes (photo, contour, thermal, score) with a drone overhead and a vibrating coffee cup
A single hairline crack, read four ways: by photograph, contour, heat, and score. The cup on the asphalt registers what the index does not.

You know the feeling well. You drive along calmly, the coffee cup settled in its holder, and then a wheel drops into a pothole you never saw. The car shudders for half a second, the cup sloshes, and you mutter a sentence not meant for children’s ears. Within minutes we forget the moment and continue down the road, as though the pothole were a passing accident that does not concern us. Yet that pothole began months earlier as a thin hairline crack, and anyone who studied it could have read its fate. What is new is that the observer is no longer a man walking the road’s edge in a reflective vest with a notebook. It is a small aircraft hovering thirty meters overhead, with an algorithm behind it that never tires. Can this new eye end the jolt itself before it happens?

The Inspector Standing at the Edge of Danger

Let us begin with the old method. The traditional survey of pavement condition depends on an inspector who walks or crawls slowly along the road and records what his eye sees. The studies reviewed in this report describe the method as subjective, exhausting, and slow, and as exposing those who perform it to the hazards of traffic. Two inspectors may rate the same crack at two different levels of severity, and a single inspector may pass two different verdicts on the same street on two different mornings.

To bring order to this chaos, engineers devised the Pavement Condition Index, adopted in the American standard ASTM D6433. Its idea is elegant. The road is divided into sample units, and each unit is recorded for the type, severity, and density of its distresses. These data are then translated into “deduct values” subtracted from one hundred. A score of one hundred means an excellent road, and zero means a road that has failed. It works like a medical checkup that condenses the body of the road into a single number, one that engineers, decision-makers, and municipal budgets can all understand.

But this standard was designed originally for manual surveys, as Yang and his colleagues note, and it does not readily accommodate automated processing. The number that looks neutral is in fact born of a tired human eye and of an inspection that takes days to cover a short stretch. The question that occupied researchers was whether the rigor of the index could be kept while the fragility of the inspector was removed. The answer began in the sky.

An Eye That Flies Above the Asphalt

The comprehensive review by Manjusha and Sunitha reveals that multi-rotor drones, such as quadcopters and hexacopters, are the most widely used platforms for collecting pavement data, because they are agile, long-ranging, and inexpensive. Data quality does not come by chance. Flight altitude, speed, the overlap between images, the ground sampling distance, and the camera angle all determine it. For three-dimensional modeling, the review recommends roughly seventy-five percent forward overlap and sixty percent side overlap. In other words, every point on the road must appear in several images so that the software can rebuild it in three dimensions.

Field examples show what this means in practice. Astor and his colleagues flew a Phantom 4 Pro at altitudes between thirteen and twenty meters and captured 1,223 photographs with seventy percent overlap. They obtained a georeferenced orthomosaic, a composite map stitched from the images, with a spatial resolution finer than one centimeter. In China, Cheng and his team flew an M300 drone at thirty meters with a high-resolution camera and captured 2,440 images of roads in the city of Tianjin, covering six types of distress.

The images then pass through photogrammetry software, which produces orthomosaics, digital surface models, and three-dimensional point clouds. These products become the visual medium on which distress is identified and its dimensions measured, so the index can be calculated without anyone visiting the site. The inspector thus moves from the street to the office, and a road once measured in footsteps is now measured in pixels. But an image alone is not a verdict, and it still needs someone to read it.

An Algorithm Learns to See Cracks

Here deep learning enters, and its most prominent family in this field is YOLO, short for “You Only Look Once.” Its premise is that a neural network scans the entire image in a single pass and locates the distresses and identifies their types, fast enough for real-time work. Liu and his colleagues developed a model derived from the eighth version of YOLO and named it MASL-YOLO. They added modules for extracting features at multiple scales, focusing attention, and widening the field of view. They tested it on 7,281 images containing transverse cracks, longitudinal cracks, alligator cracking, potholes, and cracks around manhole covers. The model reached a mean precision of 85.5 percent at 73.7 frames per second. Mounted on a vehicle equipped with a visual sensor and a positioning unit, it demonstrated real-time detection while the vehicle traveled at forty kilometers per hour.

Cheng and his colleagues built RLD-Net on the foundation of YOLOX, a model designed specifically for drone imagery. It includes modules that reduce background noise and merge fine detail with broader context. Its accuracy rose to 84.7 percent, against 78.9 percent for the baseline model.

The research did not stop with this family. Rathod and his colleagues combined vision models based on transformers with a recurrent neural network tuned by an algorithm inspired by the black widow spider, and they reached an accuracy of 97.5 percent on drone images they collected themselves. Wu and his colleagues showed that a Faster R-CNN network can detect cracks in concrete from thermal infrared images, even though the training set did not exceed two hundred images. Hu and his colleagues carried the idea to the bridge deck. They merged drone-derived maps with geographic information systems to map cracks spatially on a real bridge in Atlanta.

All these figures look impressive, but they answer only one question: did the machine see the crack? The architect’s and the planner’s question is different. How does what the machine saw become a number on which a decision can be built?

When an Image Becomes a Number Between Zero and One Hundred

The most important link in the chain connects the detection of distress to the calculation of the index. Wu and his colleagues proposed an integrated pipeline as early as 2018. It begins with collecting visible and thermal images by drone. The neural network then performs inference, and severity is determined by measuring the dimensions of each distress computationally. Deduct values follow according to the standard, and the index is finally obtained by subtracting the highest corrected deduct value from one hundred.

Field evidence came later. Yang and his team tested a system called CrackNet, which automatically detects alligator cracking and longitudinal and transverse cracks in two- and three-dimensional images. A survey vehicle, not a drone, captured those images, and the team then calculated the index automatically. At two sites in the state of Maryland, their automated results for 2018 correlated with the manual results for 2016 at a level of 90 percent at one site and 92 percent at the other.

The study closest to the spirit of this article was conducted by Astor and his colleagues on a stretch of the Bandung–Subang Highway in Indonesia, one and a half kilometers long. They divided it into sixty-nine sample units for the Pavement Condition Index and fifteen segments for an alternative measure, the Surface Distress Index, and they compared drone-derived models against manual measurements. The first index showed strong agreement, with a coefficient of determination of 86 percent, and analysis-of-variance tests found no statistically significant difference between the two methods. The alternative index proved weaker, with a coefficient of determination of 0.653. The researchers explain this by noting that the first index captures nineteen types of distress against only four for the second. A richer description of the ailment, in other words, produces a more precise diagnosis.

Yet these promising results carry within them a warning that must not be lost in the enthusiasm.

The Beautiful Numbers That May Deceive Us

When successful studies accumulate, caution becomes a virtue. Zihan and his colleagues conducted a meta-analysis of twenty-eight studies, ten on object detection and eighteen on semantic segmentation, which colors every pixel according to what it represents. The mean performance score was 80 percent for the first type and 86 percent for the second. These are excellent numbers, but the researchers paused at the source of the data. Studies that collected their images with hand-held devices recorded higher results than those that used survey vehicles equipped with moving laser scanners. This suggests that static studies may overstate a model’s readiness for real-world work.

The irony is plain: a real road does not stand still to be photographed. It contains shadows, movement, oil stains, and glare from the sun. The researchers also observed that many studies do not disclose the basic error measures adequately and do not discuss how they evaluated their errors, and they called for a standardized reporting method. They did not find image size to be an influential source of variance, which means that cutting images into small patches does not by itself weaken the reported performance.

To this must be added the absence of unified, open datasets that would allow fair comparison, despite efforts such as the RDD2022 collection, which gathered data from six countries. To ease the scarcity of data, researchers commonly resort to transfer learning from models pre-trained on general image collections. The gap between laboratory and street quickly becomes clear when we leave the glossy asphalt and meet the obstacles of reality.

Trees, Depth, and Roads Without Asphalt

The studies reveal recurring obstacles. The first is occlusion: trees, vehicles, and shadows hide distress from the drone’s lens. Astor and his colleagues recorded their lowest dimensional accuracy, 75.72 percent, for edge cracking, because trees cover the road’s edges. Their highest accuracy, 97.86 percent, came for joint reflection cracking. The researchers suggest using LiDAR sensors, which can penetrate vegetation cover. The second obstacle is environmental variability. Liu and his team turned to data augmentation to simulate rain, snow, and fog so that their model would hold up. The third is fine cracks, which grow harder to detect the higher the drone flies.

The hardest obstacle of all is depth. Three-dimensional models derived from drone imagery may fail to capture the depth of depressions and ruts accurately, because of the limited resolution of the digital elevation model. Here the whole equation stumbles, because the standard requires severity to be determined by measuring crack width, pothole depth, and rut depth. Yang and his colleagues add that current automated applications rely mainly on cracking data, while non-cracking distresses such as rutting, raveling, bleeding, and patching require further algorithm development.

The study by Khilji and his colleagues opens up a neglected dimension: unpaved roads, which remain the true arteries in many regions. The researchers built a two-stage framework. The first stage separates the road surface from its surroundings, and the second searches within it for potholes, washboarding, and ruts. The MobileNetV2 network performed best, with segmentation accuracy above 93.5 percent for the road and 86 percent for the distresses. They also noted that flying slowly, at three meters per second, gives better results than five meters per second. Speed, it seems, is always bought at the price of accuracy.

Conclusion

The drone will not prevent the next pothole simply by flying over it. It offers something harder and more valuable: to see the road before it hurts us. But between the image of a crack and the decision to repair it lies a long chain of assumptions: a standard designed for a human eye, laboratory data cleaner than reality, trees that hide the road’s edge, and a depth the lens cannot capture with precision. The deeper question may run deeper than the technology itself. If the algorithm one day becomes able to issue a number for every street in the city, will maintenance go first to the street in the worst condition or to the street with the loudest voice? And when the coffee cup trembles in your hand again at the same pothole, will you ask about the asphalt, or about who decided that its number was enough to postpone its repair?


References

Liu, Z., et al. “Real-Time Pavement Distress Detection Based on Deep Learning and Visual Sensors.” Road Materials and Pavement Design, 2024.

Manjusha, M., and Sunitha, V. “A Review of Advanced Pavement Distress Evaluation Techniques Using Unmanned Aerial Vehicles.” International Journal of Pavement Engineering, 2023.

Astor, Y., et al. “Unmanned Aerial Vehicle Implementation for Pavement Condition Survey.” Transportation Engineering, 2023.

Cheng, H., et al. “RLD-Net: An Enhanced Deep Learning Model for Accurate Pavement Distress Detection Using UAV Captured Images.” Multiscale and Multidisciplinary Modeling, Experiments and Design, 2025.

Rathod, V. V., Rana, D. P., and Mehta, R. G. “Deep Learning-Driven UAV Vision for Automated Road Crack Detection and Classification.” Nondestructive Testing and Evaluation, 2024.

Wu, W., et al. “Coupling Deep Learning and UAV for Infrastructure Condition Assessment Automation.” IEEE International Smart Cities Conference, 2018.

Hu, D., Yee, T., and Goff, D. “Automated Crack Detection and Mapping of Bridge Decks Using Deep Learning and Drones.” Journal of Civil Structural Health Monitoring, 2024.

Khilji, T. N., Loures, L. L. A., and Azar, E. R. “Distress Recognition in Unpaved Roads Using Unmanned Aerial Systems and Deep Learning Segmentation.” Journal of Computing in Civil Engineering, 2021.

Zihan, Z. U. A., Smadi, O., Tilberg, M., and Yamany, M. S. “Synthesizing the Performance of Deep Learning in Vision-Based Pavement Distress Detection.” Innovative Infrastructure Solutions, 2023.

Yang, G., et al. “Field Performance of Deep-Learning Based Fully Automated Cracking Analysis and Its Potential for PCI Surveys.” Airfield and Highway Pavements Conference, 2019.

Further Reading From ArchUp

  • |

    The Door Is the Smallest Urban Plan

    A bus door configuration serves as a foundational element of urban design, significantly influencing transit efficiency and street space. Multi-door…

  • The Urbanism of Digital Governance: How Virtual Spaces Redefine Municipal Territory and Civic Equity

    For decades, the architecture of governance manifested as monolithic concrete edifices, labyrinthine corridors winding through municipal halls, and prolonged queues…

  • Architect Salaries in 2025: The Economic Factors Reshaping Earnings

    Introduction: A Defining Moment for Architects’ Salaries As we step into 2025, the architectural profession finds itself at a crossroads,…

  • Understanding the Weight of Backfill Materials: Implications for Construction and Architecture

    Understanding the Weight of Backfill Materials: Implications for Construction and Architecture When planning a construction project, one of the most…

  • | |

    The Kingdom of the Stone Markers

    From an aircraft, or from a satellite, the oldest human decision in Arabia is still visible. Across the black lava…

Leave a Reply

Your email address will not be published. Required fields are marked *