Qianhui Li, Hua Xiang, Tongxi Wang, Wei Zhan
Industrial assembly videos contain substantial temporal redundancy, creating a need for preprocessing methods that remain usable when task-specific labels, retraining, or opaque learned selection rules are undesirable. This study designs a deterministic, training-free skeleton-semantic keyframe extraction framework for fixed-view monocular assembly video preprocessing. OpenPose BODY-25 keypoints are confidence-filtered, normalized, interpolated, and represented by eight upper-body joints. Neck-relative bilateral-wrist activity controls an inverse threshold, while sequential reference updates and a maximum-gap rule provide an explicit output budget. At 271 frames per analyzed file, the method obtains CVsd = 0.3340 ± 0.1117 and MSD = 0.7781 ± 0.1003, compared with 1.0681 ± 0.0754 and 0.2961 ± 0.0305 for pretrained MViT-V2-S feature clustering. Thus, under the shared normalized skeleton evaluation, the method reduces semantic-distance variation by 68.7% and increases adjacent pose-state separation by 162.8% relative to MViT. Matched ablation identifies the skeleton representation as the primary source of this organization; adaptive thresholding provides smaller, budget-dependent density adjustment. A second independent human annotation and adjudication establish a 34-transition event reference. Event results reveal a complementary trade-off: the proposed method retains 79.41% of transitions within ±0.5 s, whereas several temporal and appearance-based baselines retain 97.06-100%, showing that pose-state diversity and process-boundary coverage are different sampling objectives. Matched-baseline observed-joint-only and confidence-weighted analyses show that interpolation contributes to numerical regularity, but the proposed method retains substantially lower semantic-distance variation and greater adjacent pose-state separation than all evaluated baselines under both missing-data treatments. These results establish the method as an interpretable, budget-controllable preprocessing strategy for pose-oriented review and data reduction in the evaluated fixed-view setting. Validation across independently recorded production environments and task-specific downstream systems is the next step toward broader deployment.