ข้ามไปยังเนื้อหา

การส่ง String

ฟังก์ชัน WebAssembly รับและส่งกลับได้เฉพาะ numeric type เท่านั้น — integer และ float ไม่มี string type วิธีแก้ไขมาตรฐานคือ encode ข้อมูล string เป็น byte ใน linear memory แล้วส่ง integer coordinate (pointer และ length) เพื่อระบุตำแหน่งของ byte เหล่านั้น

การส่งกลับ JavaScript string ไปยัง Wasm และกลับมามี 5 ขั้นตอน:

  1. Encode string เป็น Uint8Array ของ UTF-8 byte โดยใช้ TextEncoder
  2. Write byte เหล่านั้นลงในพื้นที่ที่กำหนดของ Wasm memory
  3. Call ฟังก์ชัน Wasm พร้อม byte offset (pointer) และจำนวน byte (length)
  4. Wasm ทำงานและส่งกลับ pointer และ length ของ byte ผลลัพธ์
  5. Decode byte ผลลัพธ์กลับเป็น JavaScript string โดยใช้ TextDecoder
flowchart LR
  A["JS string\n'Hello'"] --> B["TextEncoder\n-> Uint8Array"]
  B --> C["mem.buffer\n(write at ptr)"]
  C --> D["Wasm function\n(ptr, len)"]
  D --> C
  C --> E["TextDecoder\n-> JS string"]
การส่ง string ผ่าน linear memory

TextEncoder.encode(str) แปลง JavaScript string เป็น Uint8Array ของ UTF-8 byte อักขระหลาย byte (emoji, ตัวอักษรที่มีเครื่องหมายกำกับ, ตัวอักษร CJK) ให้ byte มากกว่าจำนวน character เสมอ ดังนั้น memory operation ควรใช้ byte length เสมอ ไม่ใช่ string length

เลือก address ใน Wasm memory ที่คุณควบคุม (เช่น พื้นที่ static เริ่มต้นที่ byte 256) และเขียน byte ที่ encode ไว้:

const encoded = new TextEncoder().encode(str);
const ptr = 256;
new Uint8Array(mem.buffer).set(encoded, ptr);

หลังจากเรียกฟังก์ชัน Wasm คุณจะได้ (ptr, len) กลับมา แล้ว slice memory view และ decode:

const decoded = new TextDecoder().decode(
new Uint8Array(mem.buffer, outPtr, outLen)
);

Constructor Uint8Array แบบ 3 argument สร้าง view เริ่มต้นที่ outPtr มีจำนวน element outLen — ไม่มี copy, ไม่มีการ allocate

module Wasm ด้านล่าง expose ฟังก์ชัน echo ที่ส่งคืน (ptr, len) เดิมที่รับมา ฝั่ง JavaScript encode "Hello Wasm" เขียนที่ offset 256 เรียก echo แล้ว decode ผลลัพธ์

WebAssembly
ตัวเลือกBenefitCost
ส่ง (ptr, len) ผ่าน linear memory แทน string โดยตรงzero-copy อ่าน/เขียน byte ได้เร็ว ไม่ต้อง serialize ซ้ำผ่าน host boundaryต้อง encode/decode UTF-8 เองและคำนวณ byte length ให้ถูกต้องทุกครั้ง
กำหนด address คงที่สำหรับเขียน string (manual layout)เรียบง่าย ไม่ต้องมี allocator หรือ runtime เพิ่มไม่มี GC คอยเตือนเมื่อ string ใหม่ยาวเกินพื้นที่เดิม ต้องเผื่อขนาด buffer เอง
  • ใช้ str.length (จำนวน character) แทน encoded.length (จำนวน byte จริงหลัง UTF-8 encode) ตอนคำนวณ length ที่จะส่งให้ Wasm ทำให้ string ที่มีอักขระหลาย byte เช่น emoji หรือภาษาไทยถูก truncate หรือ decode ผิด
  • เขียน byte ที่ offset คงที่ (เช่น 256) โดยไม่เผื่อพื้นที่ให้พอกับ string ที่ยาวขึ้นเรื่อยๆ ทำให้ข้อมูลใหม่ทับ region อื่นที่ยังใช้งานอยู่
  • เก็บ typed-array view ไว้ใช้ซ้ำข้าม call แล้วลืมว่าถ้ามี memory.grow เกิดขึ้นระหว่างทาง view เก่าจะถูก detach ทำให้เขียนหรืออ่าน string ผิดตำแหน่ง

💡 ตัวอย่างจากของจริง

FFmpeg.wasm ส่ง filename และ metadata string ผ่านรูปแบบ (ptr, len) เดียวกันนี้ข้าม linear memory เพื่อสั่งงาน decoder ที่คอมไพล์มาจาก C โดยไม่ต้องแปลง string ผ่าน glue code ที่ช้าลง

วิธีส่ง JavaScript string ไปยังฟังก์ชัน WebAssembly คือ?
TextEncoder.encode(str) คืนค่าอะไร?
รูปแบบมาตรฐานของ Wasm ในการแทน string คือ?