生成式AI实用指南:使用Transformer和扩散模型(影印版)
生成式AI实用指南:使用Transformer和扩散模型(影印版)
Omar Sanseviero, Pedro Cuenca, Apolinário Passos, Jonathan Whitaker
出版时间:2025年04月
页数:396
“这本书是开发者掌握过去十年最大AI革命背后的工具与概念的必备指南。”
——Lewis Tunstall
Hugging Face机器学习工程师,Natural Language Processing with Transformers合著者
“这本书是生成式AI的最佳入门之选:从全面的解释到贴心的建议和动手练习,应有尽有。”
——Luba Elliott
AI艺术策展人,elluba.com

通过这本实用的动手指南,你可以学习如何使用生成式AI技术创造全新的文本、图像、音频,甚至音乐。你将了解最先进的生成模型的工作原理,学习如何根据需求对其进行微调和适配,以及如何组合现有的构建模块来创造新的模型和进行不同领域的创意应用。
这本入门书从理论概念着手,然后指导读者开展实际应用,并提供了大量代码示例和易懂的插图。你将学习如何使用开源库来利用transformer和扩散模型进行代码探索,并研究若干现有项目来帮助指导你的工作实践。
● 构建和自定义能够生成文本和图像的模型
● 探索使用预训练模型与微调自定义模型之间的权衡
● 创建并使用能够以任意风格生成、编辑、修改图像的模型
● 定制transformer和扩散模型以满足多种创意需求
● 训练能够反映个人独特风格的模型
  1. Preface
  2. Part I. Leveraging Open Models
  3. 1. An Introduction to Generative Media
  4. Generating Images
  5. Generating Text
  6. Generating Sound Clips
  7. Ethical and Societal Implications
  8. Where We’ve Been and Where Things Stand
  9. How Are Generative AI Models Created?
  10. Summary
  11. 2. Transformers
  12. A Language Model in Action
  13. A Transformer Block
  14. Transformer Model Genealogy
  15. The Power of Pretraining
  16. Transformers Recap
  17. Project Time: Using LMs to Generate Text
  18. Summary
  19. Exercises
  20. Challenges
  21. References
  22. 3. Compressing and Representing Information
  23. AutoEncoders
  24. Variational AutoEncoders
  25. CLIP
  26. Alternatives to CLIP
  27. Project Time: Semantic Image Search
  28. Summary
  29. Exercises
  30. Challenges
  31. References
  32. 4. Diffusion Models
  33. The Key Insight: Iterative Refinement
  34. Training a Diffusion Model
  35. In Depth: Noise Schedules
  36. In Depth: UNets and Alternatives
  37. In Depth: Diffusion Objectives
  38. Project Time: Train Your Diffusion Model
  39. Summary
  40. Exercises
  41. Challenges
  42. References
  43. 5. Stable Diffusion and Conditional Generation
  44. Adding Control: Conditional Diffusion Models
  45. Improving Efficiency: Latent Diffusion
  46. Stable Diffusion: Components in Depth
  47. Putting It All Together: Annotated Sampling Loop
  48. Open Data, Open Models
  49. Project Time: Build an Interactive ML Demo with Gradio
  50. Summary
  51. Exercises
  52. Challenge
  53. References
  54. Part II. Transfer Learning for Generative Models
  55. 6. Fine-Tuning Language Models
  56. Classifying Text
  57. Generating Text
  58. Instructions
  59. A Quick Introduction to Adapters
  60. A Light Introduction to Quantization
  61. Putting It All Together
  62. A Deeper Dive into Evaluation
  63. Project Time: Retrieval-Augmented Generation
  64. Summary
  65. Exercises
  66. Challenge
  67. References
  68. 7. Fine-Tuning Stable Diffusion
  69. Full Stable Diffusion Fine-Tuning
  70. DreamBooth
  71. Training LoRAs
  72. Giving Stable Diffusion New Capabilities
  73. Project Time: Train an SDXL DreamBooth LoRA by Yourself
  74. Summary
  75. Exercises
  76. Challenge
  77. References
  78. Part III. Going Further
  79. 8. Creative Applications of Text-to-Image Models
  80. Image to Image
  81. Inpainting
  82. Prompt Weighting and Image Editing
  83. Real Image Editing via Inversion
  84. ControlNet
  85. Image Prompting and Image Variations
  86. Project Time: Your Creative Canvas
  87. Summary
  88. Exercises
  89. References
  90. 9. Generating Audio
  91. Audio Data
  92. Speech to Text with Transformer-Based Architectures
  93. From Text to Speech to Generative Audio
  94. Evaluating Audio-Generation Systems
  95. What’s Next?
  96. Project Time: End-to-End Conversational System
  97. Summary
  98. Exercises
  99. Challenges
  100. References
  101. 10. Rapidly Advancing Areas in Generative AI
  102. Preference Optimization
  103. Long Contexts
  104. Mixture of Experts
  105. Optimizations and Quantizations
  106. Data
  107. One Model to Rule Them All
  108. Computer Vision
  109. 3D Computer Vision
  110. Video Generation
  111. Multimodality
  112. Community
  113. A. Open Source Tools
  114. B. LLM Memory Requirements
  115. C. End-to-End Retrieval-Augmented Generation
  116. Index
书名:生成式AI实用指南:使用Transformer和扩散模型(影印版)
国内出版社:东南大学出版社
出版时间:2025年04月
页数:396
书号:978-7-5766-2006-1
原版书书名:Hands-On Generative AI with Transformers and Diffusion Models
原版书出版商:O'Reilly Media
Omar Sanseviero
 
Omar Sanseviero曾任Hugging Face的首席llama官以及平台和社区负责人。他在开源、产品、研究、技术团队的交叉领域拥有丰富的经验
 
 
Pedro Cuenca
 
Pedro Cuenca是Hugging Face的机器学习工程师。
 
 
Apolinário Passos
 
Apolinário Passos是一名艺术家及Hugging Face的机器学习艺术工程师,为创意和艺术社区提供与AI模型互动的技术和工具。
 
 
Jonathan Whitaker
 
Jonathan Whitaker是一名数据科学家和深度学习研究员。除在answer.ai的研发工作外,他还通过YouTube频道DataScienceCastnet和各种免费在线资源分享知识。
 
 
The animal on the cover of Hands-On Generative AI with Transformers and Diffusion Models is the giant African swallowtail butterfly (Papilio antimachus).
The giant African swallowtail is one of the largest species of butterfly, with a wingspan of up to 9–10 inches (around the size of, say, a dinner plate or vinyl record). Yet, for all of their impressive size, relatively little is known about these colossal insects.
First discovered in 1782, swallowtails live in the tropical rainforests of west and central Africa, where they spend most of their time in the forest canopy; males have occasionally been observed mud-puddling at the forest floor, a behavior in which butterflies aggregate on wet organic matter such as soil or dung in a quest for nutrients.
Because of their diet, giant African swallowtails are highly toxic, and they have no known predators. Though there have been some reports of a population decline resulting from habitat destruction and poaching (specimens are highly prized and can fetch prices of over $1,000), the IUCN has listed the giant African swallowtail as Data Deficient: more information is needed in order for a conservation assessment to be made. Many of the animals on O’Reilly covers are endangered; all of them are important to the world.
购买选项
定价:184.00元
书号:978-7-5766-2006-1
出版社:东南大学出版社