論文ID: 2026EAP1086
The Forward-Forward (FF) algorithm replaces backpropagation with two forward passes—positive for real data and negative for synthetic inputs—optimizing local ”goodness” measures per layer. By avoiding large-scale gradient storage and repeated backward sweeps, FF significantly lowers computational overhead and energy consumption, making it attractive for low-power or edge devices. On the other hand, in the inference process of the FF algorithm, inference is attempted for each output class, and the one with the maximum output is selected. Therefore, the inference time increases proportionally with the number of output classes. Thus, we propose a binary inference method for the FF algorithm, implemented through a parallel classifier architecture. By using our proposed method, this formulation reduces the number of required forward evaluations from O(N) to O(log N) for an N-class problem. Under sufficient hardware parallelism, this property enables inference latency that is independent of the number of classes. However, this approach may fail to achieve sufficient performance as the number of classes increases. One possible cause is that the information between individual bits is not sufficiently exploited. In this paper, we further demonstrate that inference accuracy can be improved by introducing the concept of hierarchical structure, in which the first stage performs inference using multiple units. The effectiveness of the proposed method is demonstrated on the MNIST database and the EMNIST database. It should be emphasized that the proposed hierarchical inference does not rely on statistical bootstrapping or data resampling. Instead, multiple inference units are trained independently and their goodness outputs are integrated in a deterministic hierarchical manner to enforce consistency among bit-wise predictions.