基于计算机视觉双项目实践的本科生科研能力培养路径探析
摘 要:本文结合我在本科阶段参与的两个计算机视觉项目,整理科研入门时遇到的问题和处理过程。两个项目分别涉及创造性图像生成和物联网实时街景分割,我主要负责模型搭建、训练调试和结果整理。开始时,我尝试在个人电脑上完成训练,但显存不足、等待时间长等问题影响了实验进度,后来将主要训练任务转到云端GPU服务器。使用服务器后,又需要学习环境配置、数据上传和日志保存。模型能够运行以后,还要对照输入图像检查结果,再决定如何修改。文中结合保存下来的实验数据,说明分割模型不同配置的结果,以及图像生成效果与运行时间之间的取舍。这些经历让我对模型原理有了更具体的认识,也让我开始重视实验记录和结果说明。本文主要是个人实践总结,所列技术指标不用于衡量学生科研能力的高低。
关键词 :计算机视觉;创造性图像生成;物联网实时街景分割;云端GPU服务器
正文
可下载并阅读全文PDF,请按照本文版权许可使用。
Download the full text PDF for viewing and using it according to the license of this paper.
An Analysis of the Pathway for Cultivating Undergraduate Research Skills Based on Computer Vision Dual-Project Practice
Abstract: This paper describes my experience in two undergraduate computer vision projects: creative image generation and real-time street-scene segmentation for Internet of Things applications. I mainly worked on model implementation, training, debugging, and organizing results. I first tried to train the models on a personal computer. Limited GPU memory and long training times led me to move the main training jobs to rented cloud GPUs. I then had to learn how to set up the environment, upload data, and save logs. Once the models could run, I compared their outputs with the input images to decide what to change. Using saved experiment records, the paper discusses the results of different segmentation configurations and the trade-off between image generation results and running time. These tasks helped me understand the models more clearly and pay more attention to keeping records and explaining results. This is an account of my own experience; the technical metrics are not measures of students’ research skills.
Keywords: Computer Vision; Creative Image Generation; Real-time Street Scene Segmentation of the Internet of Things; Cloud-based GPU Servers
参考文献
- LeCun Y, Bengio Y, Hinton G. Deep learning[J]. Nature, 2015, 521(7553): 436-444.
- Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Advances in Neural Information Processing Systems. 2017: 5998-6008.
- Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial nets[C]//Advances in Neural Information Processing Systems. 2014: 2672-2680.
- Gatys L A, Ecker A S, Bethge M. Image style transfer using convolutional neural networks[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016: 2414-2423.
- Johnson J, Alahi A, Fei-Fei L. Perceptual losses for real-time style transfer and super-resolution[C]//European Conference on Computer Vision. Cham: Springer, 2016: 694-711.
- Huang X, Belongie S. Arbitrary style transfer in real-time with adaptive instance normalization[C]//Proceedings of the IEEE International Conference on Computer Vision. 2017: 1501-1510.
- Deng Y, Tang F, Dong W, et al. StyTr2: Image style transfer with transformers[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 11326-11336.
- Rombach R, Blattman n A, Loren z D, et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 10684-10695.
- 陈淑环,韦玉科,徐乐,等。基于深度学习的图像风格迁移研究综述 [J]. 计算机应用研究,2019, 36 (08): 2250-2255.
- Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015: 3431-3440.
- Ronneb erger O, Fischer P, Brox T. U-Net: Convolutional networks for biomedical image segmentation[C]//Medical Image Computing and Computer-Assisted Intervention. Cham: Springer, 2015: 234-241.
- Zhao H, Shi J, Qi X, et al. Pyramid scene parsing network[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 2881-2890.
- Chen L C, Papandreou G, Kokkinos I, et al. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(4): 834-848.
- Chen L C, Papandreou G, Schroff F, et al. Rethinking atrous convolution for semantic image segmentation[EB/OL]. arXiv:1706.05587, 2017.
- Chen L C, Zhu Y, Papandreou G, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation[C]//European Conference on Computer Vision. 2018: 801-818.
- Cordts M, Omran M, Ramos S, et al. The Cityscapes dataset for semantic urban scene understanding[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016: 3213-3223.
- Chollet F. Xception: Deep learning with depthwise separable convolutions[C] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 1251-1258.
- Paszke A, Gross S, Massa F, et al. PyTorch: An imperative style, high-performance deep learning library[C]//Advances in Neural Information Processing Systems. 2019: 8024-8035.
- Linn M C, Palmer E, Baranger A, et al. Undergraduate research experiences: Impacts and opportunities[J]. Science, 2015, 347(6222): 1261757.
- National Academies of Sciences, Engineering, and Medicine. Undergraduate research experiences for STEM students: Successes, challenges, and opportunities[M]. Washington, DC: The National Academies Press, 2017.

