Journal of Advanced Computational Intelligence and Intelligent Informatics
Online ISSN : 1883-8014
Print ISSN : 1343-0130
ISSN-L : 1883-8014
最新号
選択された号の論文の31件中1~31を表示しています
Regular Papers
  • Cui Hu
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 957-965
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Financial fraud in listed companies has attracted increasing academic attention due to its severe implications for investors and market stability. Traditional methods for detecting fraudulent financial statements have proven insufficient in addressing the growing complexity and volume of financial data. This study proposes a novel hybrid model combining recursive feature elimination with cross-validation and parallel random forest (RFECV-PRF) for feature selection and a genetic algorithm-optimized LightGBM (GA-LightGBM) for financial fraud detection. The RFECV-PRF method effectively evaluates feature importance and selects the optimal subset of financial indicators, while the GA-LightGBM model enhances prediction accuracy and efficiency through global optimization of hyperparameters. Empirical tests using data from five industries—energy, materials, industrials, information technology, and healthcare—demonstrated significant improvements in classification accuracy, precision, recall, and F1 scores. The proposed framework outperforms traditional machine learning approaches such as random forest and XGBoost, achieving superior results in detecting fraudulent financial activities. This research highlights the potential of integrating advanced feature selection methods with optimized machine learning models to address the challenges of financial fraud detection and improve the reliability of corporate financial reporting.

    RFECV-PRF-GA-LightGBM model framework Fullsize Image
  • Yu Shang, Shang Xue
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 966-976
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    In the era of globalization, learning English through oral means is becoming increasingly important. However, inaccurate speech recognition and imperfect feedback mechanisms have significantly hindered the improvement of the oral ability of learners. To solve this problem, this study proposes an English oral learning speech recognition and feedback system based on a multilayer improved long short-term memory network (MLSTM). The study uses the Texas Instruments and Massachusetts Institute of Technology (TIMIT) speech database and employs a hidden Markov model (HMM) and standard LSTM systems as controls to conduct a comprehensive test of the proposed model. The experimental results showed that the MLSTM model achieved high speech recognition accuracy, with an overall accuracy of 86.3% that was significantly higher than the 65.2% of HMM and 80.1% of LSTM. In terms of feedback information targeting and effectiveness, the MLSTM model scored 4.35 and 4.5, respectively, that were suggestively better than those of the comparison models. This showed that the MLSTM model could accurately recognize speech and provide learners with highly personalized and effective feedback. The research results enrich the theory of computer-assisted language learning, provide practical and effective tools for oral English learning, and promote oral English learning toward greater intelligence and efficiency.

  • Jiaxian Chen, Shuang Xu, Yuanyuan Xing
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 977-988
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    With the advancement of the intellectualization of power systems, visible light images captured by inspection robots have found extensive applications in the condition monitoring of transmission lines. However, two key challenges persist in real-world scenarios. First, critical components such as insulators and spacer dampers are frequently configured in dense and small-scale arrangements, rendering them prone to being overlooked by traditional detection methods owing to insufficient feature extraction or inadequate contextual modeling. Second, transmission corridors are often susceptible to various safety hazards, including encroaching vegetation, accumulated water, floating plastic films, and smoke from nearby fires, each posing a significant threat to operational reliability and the safe and stable operation of the power grid. Therefore, we propose a model named Mobile-RCNN that is capable of detecting both densely arranged and small-scale components as well as safety hazards. First, we constructed a safety hazard dataset through on-site photography. Second, we introduced the Q-linear-NMS post-processing method that replaced the non-maximum suppression (NMS) approach with Soft-NMS and incorporated a linear suppression control coefficient. Then, we replaced the intersection over union (IOU) evaluation method for overlapping detection boxes with the distance-IOU (DIOU) evaluation method and introduced an adjustment parameter for the degree of score suppression. Finally, to enable the backbone feature extraction network to better learn the characteristics of safety hazards and provide rich information for the subsequent detection network, we adopted MobileNetV3 as the feature extraction network for the Mobile-RCNN. The experimental results demonstrated that the proposed model detected densely arranged targets and accurately identified the locations of safety hazards, providing robust technical support for the construction of an intelligent, reliable, and all-weather operation and maintenance system for power systems.

  • Yuhong He
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 989-1003
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    This study addresses the problems of imbalanced voice sequence, insufficient training stability, and symbol-audio modality mismatch in multi-track music generation. To this end, a gradient penalty-constrained multi-track adversarial training framework and a Wasserstein distance-driven cross-modal distribution alignment mechanism are studied and designed to optimize voice coordination accuracy and auditory perception authenticity. The core innovation lies in building a dual-module collaborative optimization architecture, pioneering gradient penalty constraints to eliminate multi-track training oscillations, and proposing Wasserstein feature space mapping mechanism to eradicate cross-modal perception mismatch, establishing a theoretical paradigm and technical path for joint optimization of voice and sound effects. Experimental verification: the voice conflict rate reaches 0.60%, the gradient norm variance is 1.13×10-3, and the training stability is controlled. The cross-modal distribution distance is 0.28, achieving precise alignment. The feature alignment error of 0.17 exceeds the technical limit, and the style fidelity is 91.7% (Baroque 94% / Jazz Blues 95%) to restore artistic expression. The dynamic expressive power is 4.66 points, approaching human creativity, and the auditory similarity is 0.89, establishing perceptual authenticity. The conflict rate of voice parts in the ablation test decreases by 52%, and cross-modal mismatch compression is reduced by 49%. Parameter sensitivity analysis shows that when the hidden space dimension is 64, the structural entropy is 0.810 and the perceptual similarity is 0.87, reaching the global optimum. This model significantly improves the coordination of multi-track structures and cross-modal perception quality, providing core technical support solutions for industrial-grade artificial intelligence (AI) music creation platforms such as film and television music composition and digital composition.

  • Haitao Song, Shatie Zuo, Aoxue Yang, Yafeng Yao, Xuzhi Lai
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1004-1014
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Tunnel drilling rig is the key equipment used for exploration in underground coal mine. Because it operates for long periods in environments characterized by high humidity, intense vibration, pressure fluctuations, and unstable geological formations, various faults tend to appear frequently. If these faults are not identified in time, they may gradually worsen and ultimately result in severe accidents. Traditional fault detection methods mainly rely on manual inspection, which makes it difficult to obtain reliable information and respond effectively under complex and changing working conditions. To overcome these shortcomings, this study proposes a fault detection approach based on a sparse autoencoder. The raw signals, including pressure, speed, and feed rate, are first preprocessed. After that, a normal operating model of the drilling rig is learned through the sparse autoencoder. Faults are then detected by comparing the real-time reconstruction errors with a preset threshold. Finally, experiments based on actual drilling data are performed, and the results demonstrate the effectiveness of the proposed method.

  • Naiwen Zhang, Haitao Song, Aoxue Yang, Chengda Lu, Haipeng Fan, Hongbo ...
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1015-1024
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    An electro-hydraulic drive system is essential for the stable operation of tunnel drilling rigs in underground coal mines. However, components such as pumps, valves, and controllers inevitably experience gradual degradation under long-term and high-load conditions. Conventional monitoring approaches often rely on labeled fault data or suffer from limited interpretability, restricting their applicability in real engineering environments. To overcome these limitations, this study proposes an unsupervised degradation trend analysis method that does not use labeled samples. A sliding-window strategy was adopted to extract key statistical features. Principal component analysis was then employed to construct a unified health index, and Z-score normalization enabled the interpretable detection of abnormal tendencies in individual features. Validation on real drilling data revealed clear degradation behaviors, such as main pump leakage and control current drift, demonstrating that the proposed method offered a lightweight and interpretable solution for trend-based condition monitoring and provided practical support for the intelligent maintenance of electro-hydraulic drive systems.

  • Jinseok Woo, Chifuyu Matsumoto, Yuka Sone
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1025-1032
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    In recent years, achieving sustainable development goals has required of personalized support systems that can adapt to diverse users through natural human–system interactions. Understanding the nonverbal information of users is essential for facilitating these interactions. In this study, we focus on human gestures as a fundamental nonverbal modality that conveys user emotions and intentions. Therefore, we propose a gesture analysis system based on the relative positions of human joints and arm orientations. Using an RGB-D camera, we acquire human skeletal information and develop a gesture recognition system. Within this system, two analytical approaches are investigated: dynamic time warping-based classification method and a neural network-based classification methods. Through experimental evaluation of a small-scale dataset, we analyze the behavioral characteristics and recognition tendencies of each approach under controlled conditions. In addition, we present several cases demonstrating the effectiveness of the proposed system and discuss its applicability.

    Nonverbal analysis using RGB-D camera Fullsize Image
  • Qifu Chen, Jiaxin Cheng, Jianqi An, Jinhua She
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1033-1043
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Accurate forecasting of the hearth lining temperature in blast furnaces is essential for operational safety and efficiency, however it remains challenging owing to the complex spatiotemporal coupling and time-lag effects among process variables. To address this issue, we present a new spatiotemporal feature modeling framework that integrates gated recurrent units (GRUs) with a dual-attention mechanism to capture multi-scale temporal dependencies and dynamically assess variable importance. A convolutional neural network module is incorporated to extract localized spatial features from the time-series data, thereby enhancing the representation of the underlying metallurgical mechanisms. The validation on real industrial data showed that the proposed model achieved a root mean square error of 0.0523 and a hit rate of 94.26%, outperforming conventional long short-term memory and GRU models. This approach offers a reliable solution for intelligent health monitoring and proactive maintenance in modern data-driven ironmaking operations under highly dynamic and uncertain conditions.

  • Tao Li, Dianwei Qian, Xinlan Guo
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1044-1055
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    With the development of the economy, urban traffic congestion has become increasingly serious, causing a series of problems such as safety and environmental pollution. To address the challenge of urban traffic congestion problems, traffic signal configuration optimization based on time-varying traffic states and real-time performance is an important direction with practical engineering significance. In this study, we design a traffic signal controller that learns online without requiring a traffic model. A model-free action-dependent adaptive dynamic programming (ADP) that employs two neural networks (an action network and a critic network) to approximate the Hamilton–Jacobi–Bellman equation is employed to provide self-learning optimization ability with varying traffic states. The proposed controller employs a neuro-fuzzy system that functions as an action network to generate control decisions utilizing expertise and mitigates the stochastic exploration inefficiency inherent in conventional ADP. ADP provides reinforcement signals that indicate a reward or punishment for the neuro-fuzzy system to guide it in adjusting its parameters. Subsequently, actions associated with lower cumulative delay are reinforced. The proposed traffic signal controller can reduce the blindness of learning by using experience and the inaccuracy of the traffic model. An artificial bee colony algorithm was employed to train the ADP to meet the real-time requirement for traffic signal control. The controller can adapt to fluctuating traffic states by training continuously, and achieves a smaller average delay in the long run. The simulation results demonstrate that proposed controller achieves a reduced delay through supervised learning with an accelerated training speed.

    Adaptive model-free neuro-fuzzy system Fullsize Image
  • Runyu Ni, Hiroki Shibata, Yasufumi Takama
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1056-1073
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    This paper proposes a category-centric initialization that introduces prior knowledge for knowledge graph embedding (KGE) at a low cost. KGE is a technology that maps symbols to embeddings to utilize large-scale knowledge graphs, and it has been widely applied because of its simplicity and efficiency. However, the initialization challenge with this technology has long been overlooked. This critical issue has implications for the training cost, the stability of training, and even the final performance of the model. KGE predominantly utilizes random initialization, which overlooks the wealth of prior knowledge embedded within knowledge graphs. To counteract this, pre-training initialization has been introduced as a way to utilize the prior knowledge. While this strategy can lead to enhanced model performance and quicker convergence rates, it increases computational demands and restricts application breadth. To address these challenges, we propose a novel initialization called category-centric initialization (CCI). CCI is designed to be universally applicable across any scenario involving the training of KGE models from scratch. It utilizes the weighted sum of the category embedding and the random embedding as the initial embedding of entities. By integrating explicit category information into the random initialization, CCI effectively utilizes prior knowledge while avoiding excessive computational cost. The results of experiments demonstrate that the proposed method can effectively reduce the training cost of advanced KGE models without degrading the final performance. Additionally, the results of experiments without category information show that our method can be applied in scenarios where explicit categories are not given to entities.

    Category-centric KGE initialization Fullsize Image
  • Ruihong Wang, Weiqi Ma, Xinnan Zhang, Mengping Lin, Ni Yan
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1074-1082
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Reasonable planning of tourist routes is crucial for enhancing tourist satisfaction. This paper addresses the problem of tourist route optimization by proposing a solution that considers multiple influencing factors. A tourist route optimization model is established with the objective of maximizing tourist satisfaction, taking into account constraints such as tourist preferences, travel time, budget, and attraction opening hours. To improve the algorithm, a heuristic information mechanism and an adaptive adjustment factor are introduced to the ant colony algorithm, enhancing its global search ability and convergence speed. Using Southwest China as a case study, the results show that the proposed approach increases tourist satisfaction by 18.28% and 11.83% compared to the minimum travel budget plan and the random plan, respectively. This study provides a more efficient and accurate solution for personalized tourist route planning.

  • Xinjian Zhang, Fan Yin, Fusheng Peng, Jie Hu, Jundong Wu
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1083-1092
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    As the emphasis on energy efficiency and environmental protection grows, an intelligent control strategy for the power generation process of gas boilers has become a key approach to optimizing industrial production. This paper presents an intelligent control strategy for a 150 MW ultra-high temperature subcritical gas boiler’s power generation process. The strategy aims to address the issue of frequent manual adjustments to the mixed gas and attemperating water valves due to fuel instability. The proposed strategy achieves stable control of the power generation process and dynamic regulation of the power generation load. It does so by monitoring key operating parameters of the gas boiler in real time, designing an intelligent control strategy that uses three key process variables (main steam temperature, gas equivalent, and attemperating water valve opening degree) as controlled parameters, and implementing expert rule-based control. Operational results demonstrate that this control strategy effectively stabilizes the main steam temperature within the specified process range and enables dynamic regulation of the power generation load. With a system utilization rate exceeding 90% and a reduction in standard coal consumption from 299.8 g/kWh under manual control to 297 g/kWh under automatic control, this strategy effectively stabilizes the power generation process. It significantly improves combustion efficiency, reduces energy consumption, and mitigates environmental pollution. Thus, it has promising practical application prospects.

    Process flow diagram Fullsize Image
  • Masahiro Suzuki, Hiroki Sakaji, Kiyoshi Izumi
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1093-1101
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Pre-trained language models (PLMs) have demonstrated high performance across various tasks and domains. Among these PLMs, Mixture of Experts (MoE) models also exhibit high performance with fewer active parameters. In domain adaptation, generally, continual pre-training is performed with existing models using domain-specific corpora. However, few efforts have been made to transform and train these models into models with MoE architectures. We propose a method to construct domain-adapted MoE models from general pre-trained models that do not initially have MoE architectures. By independently training multiple experts using domain corpora and integrating them into an MoE architecture, we constructed a domain-adapted MoE model. We performed this MoE transformation in the financial domain and verified its effectiveness in financial tasks. The evaluation results indicate that our domain-adapted MoE models perform better than those without MoE architectures. Our domain-adapted MoE models achieved consistent improvements over domain-adaptive pretraining, task-adaptive pretraining, and domain- and task-adaptive pretraining, with an average improvement of approximately 0.08 in F1 across six financial benchmark tasks for encoder-based models.

    Two-stage MoE domain adaptation Fullsize Image
  • Shiyong Geng, Jintao Chen, Kuozhan Wang, Hengyi Li, Xuebin Yue, Lin Me ...
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1102-1119
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    The aging global population has significantly increased the demand for long-term care, presenting various challenges in elderly healthcare, especially in medication management. Ensuring timely and accurate medication delivery is essential to preventing medication errors, which could otherwise lead to severe health complications. This paper proposes an innovative solution to automate and enhance medication verification in nursing homes through object detection. The proposed system leverages the YOLO-GF framework, which integrates efficient computational techniques with high detection accuracy to enable real-time medicine package detection. A hybrid down sampling (HDS) module, combining max pooling, average pooling, and 2×2 stride convolution, optimizes feature map processing by reducing computational overhead while preserving key spatial information. Additionally, an enhanced multi-scale pyramid pooling (EMSPP) technique is introduced to improve multi-scale feature aggregation, enhancing the model’s ability to capture object features at various resolutions. Unlike existing detection systems that often suffer from high computational cost or poor generalization, our approach explicitly balances speed and accuracy through lightweight architectural design. Extensive experiments validate the effectiveness of the HDS and EMSPP modules in multi-scale feature learning. The YOLO-GF framework achieves a 100.00% mean average precision (mAP) on the medicine package dataset, with an inference speed of 130.16 frames per second, fully meeting the real-time monitoring requirements. Further evaluation of public datasets, including Dish20 (98.30% mAP) and Barcodes (97.34% mAP), demonstrates that YOLO-GF outperforms existing state-of-the-art models. These results highlight the superior generalization capabilities of the proposed framework. This work establishes a novel and efficient real-time solution for medicine package verification by integrating HDS and multi-scale enhancement into a unified YOLO-based framework, significantly improving accuracy and safety over manual processes in healthcare environments. The project can be found at http://www.ihpc.se.ritsumei.ac.jp/obidataset.html.

    YOLO-GF overall framework Fullsize Image
  • Chenxu Tang
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1120-1126
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    This study briefly introduces an intelligent detection algorithm for foreign objects on solar panel surfaces, as well as an intelligent cleaning robot. In the intelligent detection algorithm, the improved Retinex algorithm was used to improve low-light images, and the You Only Look Once version 5 (YOLOv5) algorithm was used to detect foreign objects on the surface. Simulation experiments were performed. The improved Retinex algorithm was compared with the traditional Retinex and histogram equalization methods. The YOLOv5 algorithm was compared with the faster region-based convolutional neural network (R-CNN) and YOLOv4 algorithms. The surface foreign object cleaning ability of the developed intelligent robot was compared with the robot that did not use the same algorithm. The results showed that the improved Retinex algorithm could increase image brightness while preserving color. The edge strength, information entropy, and locally orderless error of the improved images were 79.8±1.6, 7.5±0.7, and 813.6±2.6, respectively. The YOLOv5 algorithm could identify and locate foreign objects more accurately, with a precision of 0.987, a recall rate of 0.985, and an F-value of 0.986. It was also discovered that the intelligent robot using the proposed surface foreign object detection algorithm cleaned foreign objects on the surface of photovoltaic panels faster and better. The time consumed in one round of cleaning was 6.2±0.1 min, and the residual foreign object on the surface was 0.7%±0.1%.

  • Xu Wang, Hiroki Iwamoto, Takashi Hasuike
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1127-1137
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    The DEA-R model, which integrates data envelopment analysis (DEA) with ratio analysis, allows for the evaluation of efficiency using ratio data. One main strength of DEA is its ability to provide concrete improvement targets for inefficient decision-making units (DMUs). However, identifying these targets within the DEA-R framework is particularly challenging because of the inherent characteristics of ratio data. To address this issue, this study proposes a novel approach for identifying concrete improvement targets for inputs or outputs within the DEA-R framework. Specifically, we construct the DEA-R efficient frontier based on a unique concept and develop an approach to identify improvement targets that lie explicitly on this frontier. This ensures that all the identified targets are DEA-R efficient, thereby guaranteeing their rationality and validity. Furthermore, the constructed frontier enables the setting of flexible improvement targets under various scenarios, thereby enhancing the practicality and adaptability of the proposed approach.

  • Ziyi Huang, Xinyu Ouyang, Haigang Zhang, Jinfeng Yang
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1138-1149
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Few-shot industrial defect detection is critically challenged by poor cross-domain generalization, where models often fail to adapt from a source domain to new target domains. Existing meta-learning paradigms also face significant constraints in addressing this. Metric-based methods are prone to overfitting, while optimization-based approaches often suffer from high computational costs and training instability. To this end, we propose R2-Net, a hybrid meta-learning framework that balances performance and efficiency. Its core contribution lies in resolving the aforementioned dilemma through a synergy of optimization and inference: we employ the efficient first-order meta-optimizer Reptile to learn a high-quality set of meta-initial parameters. Building on this foundation, the model utilizes a backbone network integrated with an attention mechanism to extract high-quality features, which are then fed into a relation network for rapid and fine-grained relational defect inference. Experimental results on the MVTec AD and NEU-CLS datasets demonstrate that our framework significantly outperforms a range of strong baseline models in cross-domain few-shot tasks, validating its effectiveness.

    The R2-Net architecture, illustrating the bi-level meta-learning process Fullsize Image
  • Zhigang Peng
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1150-1162
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    English writing texts often show complex semantic layers, implicit emotional expression, and strong temporal dependence, which leads to limitations in semantic modeling and restricted feature extraction in existing methods. To address this issue, the study constructed a fine-grained sentiment recognition model that integrates bidirectional encoder representations from transformers (BERT), bidirectional gated recurrent unit (BiGRU), convolutional neural network (CNN), and an attention mechanism. BERT was used to generate context-aware semantic representations and improve overall semantic understanding of the text. BiGRU was applied to capture bidirectional temporal dependencies and describe the dynamic evolution of emotions in discourse. CNN was employed to extract phrase-level local emotional features and enhance the detection of emotion-triggering segments. The attention mechanism was introduced to highlight key emotional information and improve feature discriminability. On this basis, a gating fusion strategy was used to dynamically integrate multi-source features, and a multi-task learning framework was incorporated. The model performed emotion intensity prediction while conducting multi-class emotion classification. In this way, fine-grained sentiment modeling was achieved from both category and intensity perspectives. The results indicate that the proposed model achieves excellent performance across several datasets, with overall accuracy and F1-score remaining above 91%, ROC-AUC reaching up to 0.944, and recognition rates for all emotion categories staying above 84%, clearly outperforming existing mainstream approaches. Ablation results demonstrated that all six modules contributed to performance improvement. On the SST-2 and IMDB datasets, the complete model achieved an accuracy of 93% and an AUC of 0.97. In summary, this research offers a compact, stable, and adaptable solution for emotion modeling in English texts, with both theoretical and practical value.

  • Sho Kawai, Yotaro Fuse, Noboru Takagi, Tatsuo Motoyoshi, Hironobu Taka ...
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1163-1174
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    This study investigates human impressions of autonomous agents that refrain from action. The increasing frequency of human–agent interactions has increased the demand for cooperative human–agent behavior. In human communication, cooperative behavior is frequently regarded as refraining from action to consider the requirements or desires of others (referred to as enryo in Japanese). However, enryo behavior is rarely investigated in human–agent interaction experiments. Investigating how people react to an agent refraining from action and observing actions of the agents would advance the society toward coexistence of humans and autonomous agents. Here, we demonstrate that humans indeed perceive enryo demonstrations in refraining-from-action agents; moreover, this behavior renders an agent more anthropomorphic and likeable than a non-considerate agent. Enryo impressions are also associated with increased human proactivity, suggesting that enryo functions as a social cue that encourages initiative actions by humans. Therefore, incorporating enryo-like behavior in agents may enhance the smoothness and cooperativeness of human–agent interactions.

    Human―agent apple-catching task Fullsize Image
  • Aiqin Wang, Wenliang Li
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1175-1186
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    A scientific and reasonable student evaluation system plays a crucial guiding role in the development of higher education. However, for the current student evaluation theories and models, their operability is relatively weak. The constructed models have poor adaptability (weak robustness), lack cross-scenario research on guiding policies, and the evaluation results of the models are also relatively single. Therefore, this study proposes a hierarchical student evaluation enhanced decision-making model which is based on a zero-order fuzzy classifier as the construction unit. By introducing the administrative decision-making characteristics of the educational administration department (such as policy orientation indicators, teaching intervention suggestions, etc.), combined with a hierarchical fuzzy system and an improved ridge regression algorithm, it achieves the collaborative optimization of evaluation efficiency and interpretability. The experimental results show that the model demonstrates excellent classification performance and semantic interpretability on the degree student evaluation dataset, and can accurately predict students’ academic performance and support personalized educational decisions.

  • Qiaole Zhu, Quan Liang, Nanlin Kuang, Jinpeng Zhang
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1187-1198
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Truck scale systems are often located in complex industrial environments, and existing inspection methods based on the appearance of an entire vehicle are difficult to accurately locate trucks in space-constrained scenarios. This study proposes GSP-YOLO, a tire detection algorithm for truck scale scenarios based on an improved YOLOv8n, that assists in locating trucks on the scale by detecting tires and enhances the detection performance in industrial environments. GSP-YOLO incorporates a global-to-local spatial aggregation module into the neck structure to improve the perception of tires on small scales. A shared detail-enhanced convolutional detection head is designed in the detection head to enhance its ability to recognize complex features while reducing the computational cost. The model also replaces conventional convolution with poly-scale convolution to enhance object recognition capability and reduce feature loss. Additionally, the wise-intersection over union loss function is employed as the bounding box regression loss to suppress competition among high-quality anchors and reduce the impact of low-quality samples, thereby improving detection performance. The experimental results demonstrated that GSP-YOLO achieves an mAP50 of 89.3% on the dataset, representing a 2.0% improvement over the baseline model, and an increase of 0.9% in mAP50–95. This model significantly enhances the detection capability within truck scale systems and ensures reliable performance in complex industrial environments.

  • Zhen Cai, Qian Wang, Xiang Wang, Pei Fu
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1199-1208
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Precise orientation control of drilling tools is fundamental to directional drilling trajectory management. This study establishes a hybrid control framework for the high-accuracy regulation of inclination and azimuth. First, we derived a kinematic model characterizing the dynamic evolution of inclination and azimuth, explicitly addressing their coupled dynamics and azimuthal time-delay effects. A comprehensive downhole motion model was developed to capture system behavior. Distinct control strategies were formulated for the inclination and azimuth subsystems and integrated into a unified hybrid architecture. Stability analysis decomposes the system into individual control loops, with inclination stability ensured through conventional criteria, while azimuth stability is transformed into a solvable linear matrix inequality problem. Experimental validation demonstrates the superior robustness and engineering applicability of the method, achieving minimal attitude deviation under perturbed conditions. This control design provides a theoretically rigorous solution that bridges advanced control theory with practical well construction requirements.

    Hybrid control of drilling attitude Fullsize Image
  • Yun Wu, Ziyi Wang, Yan Du, Jieming Yang, Kai Yang
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1209-1217
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Aiming to address the problems of interest conflict between charging stations and electric vehicle (EV) owners, as well as severe load fluctuations caused by disorderly EV charging, this paper proposes a multi-objective optimal scheduling model based on an improved NSGA-III algorithm (TSM-NSGA-III). The model utilizes dynamic electricity price as a decision variable instead of a fixed time-of-use price, with optimization objectives set to maximize charging station profit, maximize EV owner satisfaction, and minimize the load peak-valley difference rate. The TSM-NSGA-III algorithm enhances the original NSGA-III through three key improvements: (1) chaotic reverse learning to improve initial population quality, (2) the sparrow search algorithm to avoid local optima, and (3) Manhattan distance to preserve population diversity and discover potential optimal solutions. Experimental results demonstrate that the proposed method achieves a 26% faster convergence and a 9.9% higher average solution quality compared to NSGA-III. Furthermore, it obtains superior Pareto frontiers with significantly better performance in both charging station revenue and user satisfaction, effectively overcoming the algorithm’s tendencies toward premature convergence and neglect of diverse optimal solutions.

    Pareto frontiers of EV scheduling solved by various algorithms Fullsize Image
  • Hongqin Tang, Yang Shen, Jianping Zhu
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1218-1230
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    The rapid advancement of digitalization is reshaping multiple aspects of firms and transforming the nature of innovation and entrepreneurship. However, a mature solution to accurately measure the digitalization maturity of enterprises is lacking. To help firms better understand their digitalization competitiveness, this study examined Chinese listed companies and, drawing on publicly available multi-dimensional heterogeneous data, employed the skyline algorithm and entropy weight method to construct a scientific measure and evaluation of digitalization maturity. To assess the reliability and practical feasibility of the measurement, this study further conducted heterogeneity analyses across industries and regions in China based on the evaluation results. The findings indicated that firms with superior digitalization maturity were predominantly concentrated in industries that were highly sensitive to digital technologies, as well as in regions characterized by stronger resource endowments and more frequent knowledge exchanges. In contrast, the digital transformation of traditional industries, such as the real estate sector, and of the regions where these industries are concentrated remained relatively weak and required further strengthening. These findings provide significant implications for both policymakers and industry stakeholders.

    Technical framework Fullsize Image
  • Hailing Bao, Ligang Cong, Xu Liu, Qingyun Liang, Rongpu Wang
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1231-1242
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    As a core component of intelligent transportation systems, vehicular ad hoc networks (VANETs) are critical to traffic efficiency and public safety. However, their security is severely threatened by Sybil attacks, whose stealthy and destructive nature complicates detection. To address this issue, this study proposes LSTM–AttnAE, a hybrid anomaly detection model that integrates LSTM, autoencoder, and attention mechanisms. Unlike conventional methods, it leverages LSTM for temporal feature extraction, an autoencoder for self-learning normal behavioral patterns, and an attention module to dynamically prioritize critical features (boosting detection sensitivity). Experiments on the VeReMi extended dataset showed that the model outperformed traditional approaches, achieving 99% precision and 98% recall. This study provides an effective technical solution for VANET security and new insights into anomaly detection in complex networks.

    Framework of the LSTM-AttnAE model Fullsize Image
  • Muxuan Liu, Ichiro Kobayashi
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1243-1257
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Systematically comparing how linguistic representations relate to brain activity has become an important topic in the field of computational neuroscience. Prior studies have mainly relied on contextual hidden states combined with linear regression, leaving open questions about the role of static input embeddings and the benefits of nonlinear mappings. In this study, we compare input embeddings and hidden states from multiple language model families (BERT, GPT-2, and LLaMA) within both encoding frameworks, which map text features to brain responses, and decoding frameworks, which reconstruct linguistic features from brain activity. We benchmarked voxel-wise ridge regression against bidirectional long short-term memory (BiLSTMs) models, using repeat-split cross-validation and explainable variance normalization on functional magnetic resonance imaging (fMRI) data from three subjects. Our analyses demonstrate that input embeddings, despite being context-invariant, remain competitive and, in some cases, outperform hidden states, while BiLSTMs provide modest but region-specific improvements over ridge regression. Fine-grained voxel-level results further revealed distinct cortical distributions of stable versus context-dependent features. Together, these findings clarify the trade-off between predictive performance and interpretability and highlight that input embeddings offer a strong and interpretable baseline for representational alignment between language models and brain activity.

    fMRI encoding flatmap Fullsize Image
  • Yuka Sone, Jinseok Woo
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1258-1264
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Recent advances in artificial intelligence and robotics have accelerated the deployment of service robots in daily environments. However, many systems still lack adaptive responsiveness to the nonverbal behaviors of users. This study proposes a gaze-based user analysis system integrated into a mixed reality (MR) smart home environment to support attentional intention-aware human–system interactions. Rather than directly estimating emotional states, the proposed approach infers attentional intentions of users based on gaze behavior as an operational proxy for emotional attunement. Using gaze data collected through HoloLens 2, we develop a machine learning model based on a long short-term memory network combined with a mixture density network to predict future gaze coordinates in a three-dimensional space. The predicted gaze information is shared with a robotic partner to enable proactive context-aware information support. The proposed system demonstrates the feasibility of leveraging gaze prediction to anticipate user focus and provide adaptive support in MR-based smart environments.

    Gaze analysis in an MR smart home system Fullsize Image
  • Yuya Takada, Yuri Murayama, Kiyoshi Izumi
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1265-1278
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    National statistical offices are exploring alternative data sources for official statistics. Such data, including point-of-sale (POS) records and mobile phone GPS logs, are originally collected for operational rather than statistical purposes. In the era of big data, the private sector accumulates vast volumes of transaction data, and leveraging such data for official statistics has become an emerging priority. However, these efforts face significant challenges, primarily due to severe selection bias stemming from such data. Prior research has shown that, even if non-representative, transaction data can produce timely statistics via density ratio estimation methods from machine learning. As a proof of concept, that study demonstrated that preliminary estimates could be generated using biased data from a Japanese private employment agency, enabling the early release of a labor market indicator otherwise delayed by up to a year. Building on this, the present study incorporates deep learning into density ratio estimation to improve accuracy. While deep learning, when applied to density ratio estimation, is often considered prone to overfitting, this study demonstrates that it can improve estimation accuracy without overfitting. Moreover, although deep learning is typically regarded as requiring extensive hyperparameter tuning, we show that it can be implemented without a significant tuning burden, supporting its practical use in the production of official statistics.

    Errors and time lags in official statistics Fullsize Image
  • Yulia Shichkina, Aleksandr P. Stepanov
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1279-1290
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    In the paper, the novel method for generating personalized magnetic resonance imaging (MRI) reports based on key phrases is introduced. In order to demonstrate the method, an application with multiple large language model (LLM) agents was written. For establishing connection between different agents, graph-based software architecture was used. The graph consists of the following nodes. (1) The data retriever node. In this node relevant reports along with the metadata are obtained from the vector database. (2) The findings generation node. This node includes an implementation of two different variations of an MRI report’s findings section generation method. The first is based on merging already existed findings parts of radiology reports that were retrieved by key phrases into the new one. The second is based on reconstructing findings section so that it contains specific paragraphs retrieved by key phrases from the vector store. (3) The impression generation node. This node implements summarization of a findings section into an impression. The performance of a findings part generation was evaluated with BLEU, ROUGE and BERTScore metrics. Experiments have shown promising results in generation of findings parts: for the first approach mean ROUGE-L score was about 0.5, BERTScore was around 0.8, for the second approach mean ROUGE-L score was approximately 0.6, BERTScore was roughly 0.9. The first approach was around three times slower than the second. During the experiments, the dataset of 908 reports gathered from seven radiologists was used. Qualitative analysis was performed with five-point Likert scale questionary and statistically analyzed by means of Mann–Whitney U test.

    Findings section generation Fullsize Image
  • Qi Xiong, Wenhui Zhang, Jingyi Yang
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1291-1305
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    This study examines the impact of the Regional Comprehensive Economic Partnership (RCEP) on China’s intra-regional trade using the synthetic control method. By combining aggregate trade analysis with structural decomposition, we find that RCEP implementation has significantly strengthened China’s trade integration with member economies, with heterogeneous effects across trade margins. These findings suggest that regional trade agreements can promote trade expansion not only through volume growth but also through structural adjustment, offering policy-relevant insights for regional economic integration.

  • Yali Niu, Jiahao An, Shihao Zou
    原稿種別: Research Paper
    2026 年30 巻4 号 p. 1306-1317
    発行日: 2026/07/20
    公開日: 2026/07/20
    ジャーナル オープンアクセス

    Scene text recognition (STR) in natural images remains highly challenging due to the large variations in character appearance across diverse real-world conditions, such as changes in font, color, layout, and background complexity—which hinder model generalization and remain insufficiently explored. To address this issue, we propose a visual prompt-guided differential learning (VPDL) framework designed to improve the generalization capability of STR models without requiring scene-specific fine-tuning. Inspired by the human ability to reference prior visual knowledge when recognizing text, VPDL introduces a set of character-level visual prompts that guide the model in perceiving appearance variations among characters. Built upon these prompts, we develop a local-to-global differential learning strategy that enhances patch-level representations and aligns global features with character cues while preserving scene-specific information. Additionally, to mitigate exposure bias in autoregressive decoding, we replace conventional label inputs with context-aware textual prompts, encouraging the decoder to better utilize textual cues embedded in image features. Extensive experiments on widely used benchmarks and real-world datasets demonstrate the effectiveness of VPDL.

    Prompt-guided STR model Fullsize Image
feedback
Top