Artificial Intelligence and Data Science
Online ISSN : 2435-9262
A study on automating and improving the accuracy of river usage surveys using multiple multimodal models
Takashi KOJIMAKeisuke YOSHIDAEmiri HOSHIHARAKazuhiko ITO
Author information
JOURNAL OPEN ACCESS

2026 Volume 7 Issue 2 Pages 148-159

Details
Abstract

Visual utilization surveys of public spaces such as riverbanks have traditionally relied on manual counting, a method that is both labor-intensive and temporally limited. With the widespread adoption of network cameras and rapid advances in AI, continuous real-time automated monitoring is now becoming feasible. This study proposes a river usage survey methodology combining a real-time object detection model (YOLO) with multiple Vision Language Models (VLMs) to automate the analysis of camera footage.

The proposed pipeline processes each image through four stages: person detection via YOLO, Optical Character Recognition (OCR) for extracting timestamp and temperature data, attribute classification to determine the gender and age group of each individual, and group-level behavior analysis to categorize activities such as basketball, walking, running, cycling, and spectating. The system also incorporates contextual area information and applies VLM-based checks to determine the presence of pickleball nets and whether parasols are open or closed.

A central finding is that no single VLM performs optimally across all tasks. Through systematic evaluation of models including Gemma3, LLaVA, Ministral-3, and Llama3.2-vision, the authors demonstrated that selectively assigning task-specific VLMs significantly improves both accuracy and processing efficiency. For OCR, Ministral-3 achieved 100% recognition. For attribute classification, llama3.2-vision and llava-llama3 showed strong gender detection rates. For behavior analysis, Gemma3 variants produced results most consistent with manual survey data.

The system was evaluated across four GPU environments, from a high-end Nvidia H200 to the compact edge-oriented Nvidia GB10. Results confirmed that even the GB10 supports near-real-time inference, suggesting that cost-effective, continuous river monitoring using a single edge device is operationally viable. This work establishes a practical foundation for scalable, automated river space utilization surveys.

Content from these authors
© 2026 Japan Society of Civil Engineers
Previous article Next article
feedback
Top