PD03.04. Computer Vision Models for Achalasia Surgery

Clicks: 2
ID: 325711
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #370 of 453 articles by views in diseases of the esophagus : official journal of the international society for diseases of the esophagus

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 453 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
Abstract Topic Benign Disease: Esophageal Surgery Background “Computer vision” refers to a set of machine learning techniques aimed at image processing. In the context of robotic surgery, these techniques have been applied in a number of contexts, including the identification of anatomical structures, with highly accurate results. Achalasia surgical treatment involves performing a myotomy in the LEE area, and precise identification of the plane of dissection and any remaining muscle fibers is imperative for its effectiveness. We aim to develop a tool to aid in the task of identifying the correct plane of dissection and to assess myotomy quality. Methods Twenty patients undergoing robotic cardiomyotomy at a tertiary center. Images captured by the da Vinci robotic platform, with informed consent. Videos were anonymized. Frames from the myotomy step were extracted. Manual annotation by a trained surgeon using the CVAT online platform. The figure below illustrates the annotation process. Two models were trained with ~5% of the total sample using different architectures to compare performance. Model 1 U-net trained on a single dataset (baseline) Architecture based on convolutional neural networks. A light weight architecture, easy to implement and with modest hardware requirements. Built in 4 down and up sampling blocks; batch normalization in each block; and ReLU activation, trained with SparseCategoricalCrossentropy loss function, adam optimizer, for 20 epochs, with a training/validation split of 0.2. The final accuracy metrics were: Intersection over Union (mIOU): 50.4%; Precision: 0.6778; Recall: 0.5859. Model 2 U-NET with a pre-trained RESNET encoder Results Model 1 was relatively effective in detecting the myotomy area, but with a poor performance in detecting muscle fibers, which are relatively less represented in this dataset. Possible solutions proposed were: increasing training data volume and quality; utilizing pre-trained models as a base, and training for more epochs. Based on these findings, we developed a second model. Model 2 utilized a pre-trained ResNet34 (residual layers with BatchNorm + ReLU,ImageNet weights), as the encoder arm of our architecture. Validation split: 0.2. Loss: 0.5*(CrossEntropy) + 0.5*(multiclass Dice), trained for 100 epochs, trained in a dataset containing miscellaneous images, not medical data. There was a marked increase in performance metrics: mIoU: 0.6711, Precision: 0.7433, Recall: 0.7953. There was higher accuracy in predicting the myotomy area, and importantly there was increased performance in detecting muscle fibers. This change in performance is attributed to the architecture, since the dataset did not differ between models. Conclusion The current paradigm focuses on refining and adapting existing models using specialized datasets. Model performance depends primarily on data quality and volume. Key challenges include the high cost of anual annotation, the limited accuracy of automated labeling, and the need for dedicated hardware for deployment. Future efforts will involve scaling training datasets and evaluating more versatile architectures, such as transformer-based and video-based models, despite increased computational requirements. We consider the early results of this model architecture promising, and will continue to scale it with the remainder of our training data, open to refining our approach with newer machine learning techniques.
Reference Key
openalex_W7203901261 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Rodrigo Nicida Garcia, Sérgio Szachnowicz, Rubens Antônio Aissar Sallum, Ulysses jr Ribeiro, Paulo Herman
Journal diseases of the esophagus : official journal of the international society for diseases of the esophagus
Year 2026
DOI
10.1093/dote/doag077.090
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.