Researcher

Microsoftfull timeMid Level
Active

Job description

Conduct foundational and frontier research in Large Language Models (LLMs), Multimodal Large Language Models, Agentic AI, and AI Systems/Infrastructure — pushing the boundaries of what is possible in model capability, efficiency, and reliability. Design, develop, and evaluate next-generation AI models and systems, from pre-training and post-training methodologies to novel architectures and scalable inference solutions. Conceive and build AI-native products, prototypes, and demos that showcase breakthrough capabilities and translate research insights into tangible user experiences. Publish influential research at top-tier venues and contribute to the broader research community through open-source releases, technical blogs, and industry engagement. Required: Bachelor's, Master's, or PhD degree in Computer Science, Software Engineering, Electrical Engineering, or a related technical field. Candidates with strong quantitative backgrounds in fundamental disciplines — such as Mathematics, Physics, or Statistics — are equally encouraged to apply. Solid foundation in mathematics (e.g., linear algebra, probability, optimization) with demonstrated analytical and problem-solving skills. Proficient programming skills in one or more of the following: Python, C/C++, or other mainstream languages. Strong self-learning ability and intellectual curiosity, with a track record of quickly mastering new domains, tools, and technologies. Professional working proficiency in English, both written and verbal, sufficient for authoring technical papers, documentation, and cross-team communication. Research experience in one or more of the following areas: Large Language Models, Natural Language Processing, Computer Vision, Speech/Audio Processing, Multimodal AI, Reinforcement Learning, or AI Systems/Infrastructure. Publication track record at top-tier conferences or journals (e.g., ICLR, NeurIPS, ICML, ACL, EMNLP, CVPR, ICCV). Hands-on experience with large-scale distributed training, model optimization, or building end-to-end ML pipelines. Familiarity with modern deep learning frameworks (e.g., PyTorch) and large-scale computing environments. Experience shipping AI-powered features or products, demonstrating the ability to bridge the gap between research and production.