<feed xmlns="http://www.w3.org/2005/Atom"> <id>https://heyyjunn.github.io/</id><title>Min-jun, Kim</title><subtitle>데이터사이언스와 인공지능을 공부하며, Machine Learning, Deep Learning, NLP, Multimodal AI, 논문 리뷰와 프로젝트를 기록하는 기술 블로그입니다.</subtitle> <updated>2026-10-08T19:25:55+09:00</updated> <author> <name>김민준 (Min-jun, Kim)</name> <uri>https://heyyjunn.github.io/</uri> </author><link rel="self" type="application/atom+xml" href="https://heyyjunn.github.io/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="https://heyyjunn.github.io/"/> <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator> <rights> © 2026 김민준 (Min-jun, Kim) </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>[Paper] You Only Look Once: Unified, Real-Time Object Detection</title><link href="https://heyyjunn.github.io/posts/Paper-You-Only-Look-OnceUnified-Real-Time-Object-Detection/" rel="alternate" type="text/html" title="[Paper] You Only Look Once: Unified, Real-Time Object Detection" /><published>2026-10-07T02:10:22+09:00</published> <updated>2026-10-07T06:36:45+09:00</updated> <id>https://heyyjunn.github.io/posts/Paper-You-Only-Look-OnceUnified-Real-Time-Object-Detection/</id> <content type="text/html" src="https://heyyjunn.github.io/posts/Paper-You-Only-Look-OnceUnified-Real-Time-Object-Detection/" /> <author> <name>김민준 (Min-jun, Kim)</name> </author> <category term="Paper" /> <summary>(Background) What is Object Detection? object classifcation: 이미지 내 single object, output: class probability (Object class) object localization: 이미지 내 single object, output: (x,y,w,h) (Object Class, Bounding Box: 물체의 위치) object detection: 이미지 내 multiple object, output: class probabilities + (x,y,w,h) (Objet class, Bounding Box) (Background) One-Stage Detector VS. Two-Stage Detector image so...</summary> </entry> <entry><title>[Paper] Vision Transformer: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale</title><link href="https://heyyjunn.github.io/posts/Paper-Vision-Transformer-AN-IMAGE-IS-WORTH-16X16-WORDS-TRANSFORMERS-FOR-IMAGE-RECOGNITION-AT-SCALE/" rel="alternate" type="text/html" title="[Paper] Vision Transformer: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" /><published>2026-10-01T18:24:08+09:00</published> <updated>2026-10-06T23:03:33+09:00</updated> <id>https://heyyjunn.github.io/posts/Paper-Vision-Transformer-AN-IMAGE-IS-WORTH-16X16-WORDS-TRANSFORMERS-FOR-IMAGE-RECOGNITION-AT-SCALE/</id> <content type="text/html" src="https://heyyjunn.github.io/posts/Paper-Vision-Transformer-AN-IMAGE-IS-WORTH-16X16-WORDS-TRANSFORMERS-FOR-IMAGE-RECOGNITION-AT-SCALE/" /> <author> <name>김민준 (Min-jun, Kim)</name> </author> <category term="Paper" /> <summary>1. Abstract 이미지를 고정 크기의 patch로 나누고, 각 patch를 token처럼 취급하여 standard Transformer를 이미지 분류에 직접 적용함. ViT는 중간 규모의 데이터에서는 CNN보다 성능이 떨어질 수 있지만, 대규모 데이터셋으로 pre-training한 뒤 downstream task에 transfer하면 기존 CNN 기반 모델과 경쟁하거나 더 좋은 성능을 보임. 핵심적으로 CNN 없이도 Transformer만으로 image recognition이 가능하며, 데이터와 모델 규모가 커질수록 좋은 scalability를 보임. 2. Introduction Transformer는 NLP에서 대규모 pre-training과 transfer learning을 통해 큰 ...</summary> </entry> <entry><title>[Multimodal] Hugging Face Community Computer Vision Course - Multimodal Tasks and Models</title><link href="https://heyyjunn.github.io/posts/Multimodal-Hugging-Face-Community-Computer-Vision-Course-Multimodal-Tasks-and-Models/" rel="alternate" type="text/html" title="[Multimodal] Hugging Face Community Computer Vision Course - Multimodal Tasks and Models" /><published>2026-09-17T04:12:03+09:00</published> <updated>2026-10-04T09:44:06+09:00</updated> <id>https://heyyjunn.github.io/posts/Multimodal-Hugging-Face-Community-Computer-Vision-Course-Multimodal-Tasks-and-Models/</id> <content type="text/html" src="https://heyyjunn.github.io/posts/Multimodal-Hugging-Face-Community-Computer-Vision-Course-Multimodal-Tasks-and-Models/" /> <author> <name>김민준 (Min-jun, Kim)</name> </author> <category term="Multimodal" /> <summary>🔗 Reference - Community Computer Vision Course documentation, Multimodal Tasks and Models 1. Multimodal Tasks and Models 인간의 세계는 다양한 감각 입력의 조합으로 이루어져 있다. 우리는 시각, 청각, 촉각 등을 통해 세상을 인식하고 이해함. 이러한 멀티모달성(multimodality)은 여러 종류의 정보를 함께 이용한다는 점에서 하나의 종류의 정보만 처리하는 기존 단일모달 AI 모델과 구분됨. 멀티모달 모델은 텍스트, 이미지, 오디오, 센서 데이터와 같은 여러 출처의 정보를 통합함으로써 이러한 차이를 줄이는 것을 목표로 함. modality 모델에게 들어오는 정보의 종류 텍스트 → text modal...</summary> </entry> <entry><title>[Paper] AlexNet: ImageNet Classification with Deep CNN</title><link href="https://heyyjunn.github.io/posts/CV-AlexNet-ImageNet-Classification-with-Deep-CNN/" rel="alternate" type="text/html" title="[Paper] AlexNet: ImageNet Classification with Deep CNN" /><published>2026-09-14T03:27:08+09:00</published> <updated>2026-10-04T07:01:56+09:00</updated> <id>https://heyyjunn.github.io/posts/CV-AlexNet-ImageNet-Classification-with-Deep-CNN/</id> <content type="text/html" src="https://heyyjunn.github.io/posts/CV-AlexNet-ImageNet-Classification-with-Deep-CNN/" /> <author> <name>김민준 (Min-jun, Kim)</name> </author> <category term="Paper" /> <summary>Introduction 초기 객체 인식 분야는 머신러닝 기법에 주로 의존하였고, 성능 향상을 위해서 더 많은 양의 데이터셋, 더 깊은 모델, 과적합을 방지하기 위한 고도화된 학습 기법이 필요하게 되었음. CNN은 FFN보다 더 적은 파라미터와 복잡성으로 학습하기 쉬움에도 불구하고 성능차이가 적다는 장점이 있지만, 고해상도와 같은 이미지에 대해서는 비용이 많이 들어감. 기존의 얕은 CNN 으로는 큰 모델을 학습하기에는 충분하지 않았기에, 더 깊은 구조를 가진 모델을 필요로 함. 그 당시 GPU의 발전으로 인해 최적화된 2d convlution 연산을 병렬 처리할 수 있게되어 깊은 모델을 구현할 수 있게 됨. ReLU activation function, GPU 병렬처리, Drop-out, 데이터 ...</summary> </entry> <entry><title>[Paper] GPT-1: Improving Language Understanding by Generative Pre-Training</title><link href="https://heyyjunn.github.io/posts/Paper-GPT-1-Improving-Language-Understanding-by-Generative-Pre-Training/" rel="alternate" type="text/html" title="[Paper] GPT-1: Improving Language Understanding by Generative Pre-Training" /><published>2026-09-10T16:51:11+09:00</published> <updated>2026-10-08T19:20:24+09:00</updated> <id>https://heyyjunn.github.io/posts/Paper-GPT-1-Improving-Language-Understanding-by-Generative-Pre-Training/</id> <content type="text/html" src="https://heyyjunn.github.io/posts/Paper-GPT-1-Improving-Language-Understanding-by-Generative-Pre-Training/" /> <author> <name>김민준 (Min-jun, Kim)</name> </author> <category term="Paper" /> <summary>⭐️ 본문 내부의 슬라이드 이미지는 자체 제작한 프레젠테이션을 활용하였습니다. GPT-1은 대규모 비지도 텍스트로 Transformer 언어모델을 먼저 사전학습(pre-training)한 뒤, 적은 양의 라벨 데이터로 각 NLP task에 fine-tuning하면 다양한 작업에 잘 전이될 수 있다는 것을 보여준 논문 Input: 자연어 문장을 토큰 시퀀스로 변환한 것. 예: ["The", "cat", "sat", ...] Output: Pre-training 때: 각 위치에서 다음 토큰의 확률분포를 예측 Fine-tuning 때: 마지막 hidden representation을 이용해 해당 task의 정답을 출력함. 예: NLI의 entailment /...</summary> </entry> </feed>
