MOSP / 期刊 / 产教融合研究与实践 / 2026年·第4卷·第4期
研究文章

基于计算机视觉双项目实践的本科生科研能力培养路径探析

作者:卜彤*
西安工业大学 陕西 西安
*标记署名为通讯作者
提交:2026-05-31修改:2026-06-09录用:2026-07-21出版:2026-08-28
产教融合研究与实践 2026, 4(4), 1-9; https://doi.org/10.58244/rpiei.264005
基金信息:西安工业大学国家级大学生创新创业训练项目(202510702030)、西安工业大学校级大学生创新创业训练项目(X202510702184)

摘 要:本文结合我在本科阶段参与的两个计算机视觉项目,整理科研入门时遇到的问题和处理过程。两个项目分别涉及创造性图像生成和物联网实时街景分割,我主要负责模型搭建、训练调试和结果整理。开始时,我尝试在个人电脑上完成训练,但显存不足、等待时间长等问题影响了实验进度,后来将主要训练任务转到云端GPU服务器。使用服务器后,又需要学习环境配置、数据上传和日志保存。模型能够运行以后,还要对照输入图像检查结果,再决定如何修改。文中结合保存下来的实验数据,说明分割模型不同配置的结果,以及图像生成效果与运行时间之间的取舍。这些经历让我对模型原理有了更具体的认识,也让我开始重视实验记录和结果说明。本文主要是个人实践总结,所列技术指标不用于衡量学生科研能力的高低。

关键词 :计算机视觉;创造性图像生成;物联网实时街景分割;云端GPU服务器


正文

可下载并阅读全文PDF,请按照本文版权许可使用。
Download the full text PDF for viewing and using it according to the license of this paper.


An Analysis of the Pathway for Cultivating Undergraduate Research Skills Based on Computer Vision Dual-Project Practice

Abstract: This paper describes my experience in two undergraduate computer vision projects: creative image generation and real-time street-scene segmentation for Internet of Things applications. I mainly worked on model implementation, training, debugging, and organizing results. I first tried to train the models on a personal computer. Limited GPU memory and long training times led me to move the main training jobs to rented cloud GPUs. I then had to learn how to set up the environment, upload data, and save logs. Once the models could run, I compared their outputs with the input images to decide what to change. Using saved experiment records, the paper discusses the results of different segmentation configurations and the trade-off between image generation results and running time. These tasks helped me understand the models more clearly and pay more attention to keeping records and explaining results. This is an account of my own experience; the technical metrics are not measures of students’ research skills.

Keywords: Computer Vision; Creative Image Generation; Real-time Street Scene Segmentation of the Internet of Things; Cloud-based GPU Servers


 

参考文献

  1. LeCun Y, Bengio Y, Hinton G. Deep learning[J]. Nature, 2015, 521(7553): 436-444.
  2. Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Advances in Neural Information Processing Systems. 2017: 5998-6008.
  3. Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial nets[C]//Advances in Neural Information Processing Systems. 2014: 2672-2680.
  4. Gatys L A, Ecker A S, Bethge M. Image style transfer using convolutional neural networks[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016: 2414-2423.
  5. Johnson J, Alahi A, Fei-Fei L. Perceptual losses for real-time style transfer and super-resolution[C]//European Conference on Computer Vision. Cham: Springer, 2016: 694-711.
  6. Huang X, Belongie S. Arbitrary style transfer in real-time with adaptive instance normalization[C]//Proceedings of the IEEE International Conference on Computer Vision. 2017: 1501-1510.
  7. Deng Y, Tang F, Dong W, et al. StyTr2: Image style transfer with transformers[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 11326-11336.
  8. Rombach R, Blattman n A, Loren z D, et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 10684-10695.
  9. 陈淑环,韦玉科,徐乐,等。基于深度学习的图像风格迁移研究综述 [J]. 计算机应用研究,2019, 36 (08): 2250-2255.
  10. Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015: 3431-3440.
  11. Ronneb erger O, Fischer P, Brox T. U-Net: Convolutional networks for biomedical image segmentation[C]//Medical Image Computing and Computer-Assisted Intervention. Cham: Springer, 2015: 234-241.
  12. Zhao H, Shi J, Qi X, et al. Pyramid scene parsing network[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 2881-2890.
  13. Chen L C, Papandreou G, Kokkinos I, et al. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(4): 834-848.
  14. Chen L C, Papandreou G, Schroff F, et al. Rethinking atrous convolution for semantic image segmentation[EB/OL]. arXiv:1706.05587, 2017.
  15. Chen L C, Zhu Y, Papandreou G, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation[C]//European Conference on Computer Vision. 2018: 801-818.
  16. Cordts M, Omran M, Ramos S, et al. The Cityscapes dataset for semantic urban scene understanding[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016: 3213-3223.
  17. Chollet F. Xception: Deep learning with depthwise separable convolutions[C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 1251-1258.
  18. Paszke A, Gross S, Massa F, et al. PyTorch: An imperative style, high-performance deep learning library[C]//Advances in Neural Information Processing Systems. 2019: 8024-8035.
  19. Linn M C, Palmer E, Baranger A, et al. Undergraduate research experiences: Impacts and opportunities[J]. Science, 2015, 347(6222): 1261757.
  20. National Academies of Sciences, Engineering, and Medicine. Undergraduate research experiences for STEM students: Successes, challenges, and opportunities[M]. Washington, DC: The National Academies Press, 2017.

CC BY 4.0
© 2026 作者版权所有。许可出版:澳门科学出版社(MOSP),中国澳门。本文为开放获取文章,依据知识共享署名(CC BY 4.0)许可协议的条款和条件分发。 © 2026 by the authors. Published by Macao Scientific Publishers (MOSP), Macao, China. This is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY 4.0) license.
免责声明: 本期刊所发表文章中的所有陈述、观点和数据均为作者和贡献者的个人观点,不代表澳门科学出版社及/或编辑的立场。澳门科学出版社及/或编辑不对文中提及的任何思想、方法、说明或产品所导致的人身或财产伤害承担责任。 Publisher's Note: The statements, opinions and data contained in this journal are solely those of the individual authors and contributors and not of MOSP and/or the editors. MOSP and/or the editors disclaim responsibility for any injury to persons or property resulting from any ideas, methods, instructions or products referred to in the content.
Submit Your Manuscript Now