開發紀錄:幾軸、直驅還是減速、要不要動鏡——四個必須最先做、文獻卻最少的決定

這是 LocalPapa Notes 開發紀錄系列的第二十九篇,也是雲台系列的第十二篇。

上一篇把文獻攤成地圖之後,有一格特別刺眼:架構選擇的論文最少,但它是設計流程裡最先要做的決定。

原因不難理解——「兩軸還是三軸」「直驅還是減速」這類問題,答案取決於任務需求,寫不成一篇有普遍性的論文。於是它變成廠商白皮書與口耳相傳的地帶。

它們其實可以算。這篇把四個決策各自算出一個數字。

文獻基礎

架構這一塊學術文獻稀少,但不是沒有:

Hilkert, J. M. (2004). A comparison of inertial line-of-sight stabilization techniques using mirrors. Proc. SPIE 5430, Acquisition, Tracking, and Pointing XVIII.

鏡面穩定的關鍵一篇。它指出反射定律在垂直於 LOS 的軸上帶來一個固有的 2:1 關係,而且——這是最重要的一句——單純把鏡面對慣性空間穩住並不會穩住 LOS,那個 2:1 反而讓鏡面穩定架構對基座運動特別敏感。它比較了幾種結合機構與慣性/相對運動感測器的作法,逐一討論各自的取捨。下面第五節整節在展開這句話。

Dynamics Modeling and Theoretical Study of the Two-Axis Four-Gimbal Coarse–Fine Composite UAV Electro-Optical Pod (2020). Applied Sciences, 10(6), 1923.

粗精複合軸的完整處理。用有限元分析與關鍵零件應力分析做結構、框架採 7075 鋁合金以達成全機 1 kg 以下的超輕量要求;依歐拉剛體動力學模型導出兩軸四框之間的傳動路徑與運動學耦合補償矩陣。致動採超音波馬達(USM)當粗級、音圈馬達(VCM)當精級。同組另一篇把該模型代進 DOB 抑制控制,報稱擾動抑制比 PID 好上 90%、比傳統 DOB 好 25%。

Ekstrand, B. (2001). IEEE TAES 37(3). 前面幾篇一直在引的那篇——它的慣量對稱條件說明部分交叉耦合可以在架構階段設計掉,這是「架構決策先於控制設計」最直接的學理依據。

這幾篇解的是各自架構的建模與控制。沒有一篇回答「我該選哪一種」——那需要把不同架構放在同一組規格下比較,而論文通常只做一種。下面的比較是我自己算的,用的是本系列一路在用的那組常數。

本文的引用範圍

以上論文的方法與貢獻描述來自摘要與公開說明頁;本環境對多數出版商回 403,我沒有逐篇核對全文,因此沒有引用任何未經查證的實驗數值。傳動背隙的量級(諧波 1–3 arcmin、行星 3–10 arcmin)取自傳動元件的通用規格範圍,屬於業界公開的量級而非特定廠商數據;換算與比較是我自己算的。

一、四個決策,順序不能顛倒

四個架構決策,順序不能顛倒 四個架構決策,順序不能顛倒 幾軸? 決定禁區與框架數 由任務決定:可接受多大的天底禁區、需不需要影像去旋轉 直驅還是減速? 決定背隙與有效 Kt 由穩定精度規格決定:背隙一旦大於 θ_max,控制器補不回來 要不要粗精複合? 決定頻寬上限 由頻寬需求決定:單級撐不到的頻寬才需要第二級 動框架還是動鏡? 決定慣量與基座敏感度 由慣量與體積決定:但要付 2:1 的代價 順序不能顛倒的理由:軸數決定框架配置,框架配置決定慣量,慣量才輪得到談傳動與致動。先選好馬達再回頭決定幾軸,等於把前面的計算重做一次。
圖 1:四個架構決策的順序,以及每一個「由什麼決定」與「決定了什麼」。
  • 幾軸? —— 由任務決定:可接受多大的天底禁區、需不需要影像去旋轉。
  • 直驅還是減速? —— 由穩定精度規格決定:背隙一旦大於 θ_max,控制器補不回來。
  • 要不要粗精複合? —— 由頻寬需求決定:單級撐不到的頻寬才需要第二級。
  • 動框架還是動鏡? —— 由慣量與體積決定,但要付 2:1 的代價。

順序不能顛倒的理由很實際:軸數決定框架配置,框架配置決定慣量,慣量才輪得到談傳動與致動。 先選好馬達再回頭決定幾軸,等於把第二十三二十四篇那整套計算重做一次。

二、幾軸:第三軸不解 gimbal lock

幾軸各自解掉什麼——第三軸不解 gimbal lock 幾軸各自解掉什麼——第三軸不解 gimbal lock 雙軸(方位+俯仰) 三軸(+繞光軸 roll) 兩軸四框(粗精複合) 指向自由度 2(足夠指任何方向) 2(第三軸不改變指向) 2 gimbal lock 天底附近發散 依然發散——roll 軸繞的是光軸 外框可接手,不發散 影像旋轉 不補償 補償(第三軸的真正用途) 看配置 擾動抑制頻寬 受單級機構限制 同左 精級可高一個數量級 重量代價 基準 +一組馬達與框架 +精級致動器與光學 常見誤解:以為加第三軸就能繞開 gimbal lock。繞光軸的 roll 軸不改變視線方向,所以不解奇異點;要解奇異點得在外側再加一軸,那就是兩軸四框。 第三軸買到的是影像去旋轉與酬載姿態,不是指向能力。如果你的痛點是天底追不動,加第三軸不會有幫助。
圖 2:三種軸數配置各自解掉什麼。中間那欄標成橘色,因為它是最常被誤解的一個。

這是本篇最想更正的一個誤解。

繞光軸的 roll 軸不改變視線指向。 它轉的是影像的方位,不是 LOS 的方向。所以:

  • 不解 gimbal lock。第二十四篇算的天底禁區在三軸雲台上一樣存在,因為禁區來自方位軸的 1/cos(el) 發散,而第三軸完全不參與那個換算。
  • 解的是影像旋轉。載體滾轉時畫面會跟著轉,第三軸把它轉回來。對觀測與目標判讀有實質價值,但那是另一個需求。

真正要解 gimbal lock,得在外側再加一軸——那就是兩軸四框的配置。多出來的外框在內框接近奇異點時接手,讓內框始終工作在遠離 cos(el) → 0 的區域。

所以:如果你的痛點是「天底附近追不動」,加第三軸不會有幫助。 這個判斷很值錢,因為第三軸的代價是一整組馬達、框架、走線與功耗。

三、直驅還是減速:背隙是硬牆

傳動:背隙是硬牆,控制器補不回來 傳動:背隙是硬牆,控制器補不回來 0 0.5 1 2 3 角度誤差(mrad) θ_max = 0.1 mrad(本系列的穩定精度規格) 直驅(零背隙) 0 諧波減速 1–3 arcmin = θ_max 的 2.9–8.7 倍 行星減速 3–10 arcmin = θ_max 的 8.7–29.1 倍 背隙是機構的「死區」,不是可以靠回授壓下去的擾動——換向的瞬間回授根本不知道發生了什麼。諧波減速的背隙就已經是 θ_max 的 2.9–8.7 倍,所以要 0.1 mrad 等級的穩定精度,直驅幾乎是唯一選項。
圖 3:三種傳動的背隙區間,跟本系列的 θ_max 畫在同一把尺上。綠色虛線就是 θ_max。

先把單位換算清楚。1 arcmin = π/180/60 rad = 0.2909 mrad

  • 諧波減速 背隙 1–3 arcmin = 0.291–0.873 mrad
  • 行星減速 背隙 3–10 arcmin = 0.873–2.909 mrad
  • 直驅 背隙 = 0(轉子與負載軸直接耦合,力矩經磁場傳遞)

本系列一路在用的穩定精度規格是 θ_max = 0.1 mrad。對照一下:

諧波減速的背隙就已經是 θ_max 的 2.9 到 8.7 倍;行星減速是 8.7 到 29.1 倍。

這個比較之所以決定性,是因為背隙不是可以靠回授壓下去的擾動,它是機構的死區。換向的那一瞬間,馬達轉了而負載沒動,回授根本不知道發生了什麼——沒有任何控制律能補一個「感測不到」的東西。

結論很硬:要 0.1 mrad 等級的穩定精度,直驅幾乎是唯一選項。 反過來說,如果你的規格是 1 mrad 而不是 0.1,諧波減速就重新回到桌上。先確定規格,再決定傳動,順序不能反。

另外要注意諧波減速的一個內在矛盾:讓它幾乎無背隙的柔輪,本身就是會變形的零件。負載一上去它會扭轉「捲繞」——零背隙與高扭轉剛度在諧波減速上是互斥的,而扭轉剛度直接影響 f_res,也就是第二十五篇那道速率環天花板。

四、減速比並沒有讓你逃過天底禁區

天底禁區只看「有效 Kt」,跟減速比無關 天底禁區只看「有效 Kt」,跟減速比無關 0 4 8 13 17 有效 Kt = 減速比 × 馬達 Kt(N·m/A) 天底禁區半角(度) N=1 Kt_eff=0.05 禁區 0.27° N=5 Kt_eff=0.25 禁區 1.34° N=20 Kt_eff=1.00 禁區 5.36° N=50 Kt_eff=2.50 禁區 13.52° ω_max(負載端) = 0.7·V_bus / (N · Kt_motor) —— 減速比同時放大力矩、除掉轉速,兩者恰好抵消,所以禁區公式裡只剩有效 Kt。 減速沒有讓你逃過禁區,它只是換一個地方付錢:買到力矩容量,付出的是禁區變大,而且還要加上背隙。這就是為什麼高精度雲台幾乎都走直驅——不是因為直驅比較好,是因為另一條路要付兩次。
圖 4:天底禁區半角對有效 Kt 的曲線,以及四個減速比的落點。它們全都落在同一條曲線上——這就是重點。

這一節是我算出來覺得最漂亮的一段。

第二十四篇推過天底禁區:

`` 禁區半角 = asin( ω_LOS · Kt / (0.7 · V_bus) ) ``

直覺上會想:加個減速比 N,力矩放大 N 倍,是不是就能用小 Kt 的馬達、把禁區壓小?

不行。 把減速算進去:

`` 負載端有效 Kt = N · Kt_motor 負載端 ω_max = (0.7 · V_bus / Kt_motor) / N = 0.7 · V_bus / (N · Kt_motor) ``

減速比同時放大力矩、除掉轉速,兩者恰好抵消。代進禁區公式,N 完全消失——只剩下有效 Kt。

Kt_motor = 0.05、24 V 匯流排、ω_LOS = 90°/s 算幾個點:

  • N = 1(直驅)—— 有效 Kt 0.05、負載端 ω_max 19251 °/s、禁區 0.27°
  • N = 5 —— 有效 Kt 0.25、ω_max 3850 °/s、禁區 1.34°
  • N = 20 —— 有效 Kt 1.00、ω_max 963 °/s、禁區 5.36°
  • N = 50 —— 有效 Kt 2.50、ω_max 385 °/s、禁區 13.52°

四個點全部落在同一條曲線上。減速沒有讓你逃過禁區,它只是換一個地方付錢: 買到的是力矩容量,付出的是禁區變大,而且還要加上背隙。

這就是為什麼高精度雲台幾乎都走直驅——不是因為直驅比較好,是因為另一條路要付兩次。

五、粗精複合:為什麼精級的力矩需求反而更小

單級機構撐不到需要的頻寬時,加第二級。直覺上會擔心:精級頻寬高十倍,α = (2πf)² · θ_max 是平方關係,力矩需求豈不是爆炸?

算一下就知道不會。用第二十四篇的內框慣量 J = 0.0036f_bw = 12 Hz 當粗級,精級假設慣量小三個數量級(音圈馬達驅動的小鏡 vs 整個內框)、頻寬高一個數量級:

  • 粗級 —— J = 3.6×10⁻³f = 12 Hzα = 0.568 rad/s²T = 2.05×10⁻³ N·m
  • 精級 —— J = 3.6×10⁻⁶f = 120 Hzα = 56.85 rad/s²T = 2.05×10⁻⁴ N·m

α 大了 100 倍,但慣量小了 1000 倍,淨結果是精級的力矩需求只有粗級的 10%

這就是粗精複合能成立的物理理由:平方項雖然可怕,但慣量是線性項而且你可以把它做到很小。 上面那篇 Applied Sciences 用超音波馬達當粗級、音圈馬達當精級,正是這個分工——粗級管行程,精級管頻寬。

代價是:多一套致動器、多一組光學、多一層運動學耦合要補償(那篇論文的耦合補償矩陣就是在處理這個),還有整機體積與重量。單級撐得住就不要上第二級。

六、動框架還是動鏡:2:1 與一個陷阱

鏡面穩定:2:1 放大,以及一個反直覺的後果 鏡面穩定:2:1 放大,以及一個反直覺的後果 (甲)鏡面轉 θ,視線轉 2θ ΔLOS 鏡面 入射 鏡轉 15° → ΔLOS = 30°(兩倍) (乙)把鏡面對慣性空間穩住=完全沒有穩定 ΔLOS 鏡面對慣性靜止(θ = 0) 基座轉 φ 鏡不動、基座轉 20° → ΔLOS = -20°(全部漏過去) ΔLOS = 2·θ_鏡 − φ_基座 要穩住 LOS:θ_鏡 = φ_基座 / 2 Hilkert 2004 講的就是這件事:反射定律的 2:1 讓「鏡面穩定」這個說法本身就是陷阱。需求是 LOS 靜止,不是鏡面靜止——而後者反而讓基座運動 100% 通過。好處仍在:動鏡的慣量遠小於動整個框架,同樣的 LOS 角速率只需一半的機構角速率。
圖 5:反射定律的兩個後果。左格是 2:1 放大;右格是那個陷阱——鏡面對慣性空間靜止,基座運動反而 100% 漏到視線上。

用反射定律(二維,數學慣例):出射方向角 = 2α + 180° − ψ,其中 α 是鏡面法線角、ψ 是入射方向角。對兩個變數各微分一次:

`` ΔLOS = 2 · θ_鏡 − φ_基座 ``

好處在第一項。 鏡面轉 θ,LOS 轉 ——同樣的 LOS 角速率只需要一半的機構角速率。再加上動鏡的慣量遠小於動整個框架,這是鏡面穩定真正的優勢。

陷阱在第二項。 假設你「把鏡面對慣性空間穩住」,也就是 θ_鏡 = 0。那麼:

`` ΔLOS = 2 · 0 − φ_基座 = −φ_基座 ``

基座運動 100% 漏到 LOS 上,一點都沒穩到。 這正是 Hilkert 2004 那句「單純把鏡面對慣性空間穩住並不會穩住 LOS」的意思。

要真的穩住 LOS,需要的是:

`` θ_鏡 = φ_基座 / 2 ``

鏡面必須以基座角速率的一半跟著動,而不是靜止不動。 這件事違反直覺到值得寫在牆上——「穩定」這個詞在鏡面架構下指的是「穩住 LOS」,而不是「穩住鏡面」,兩者的控制目標完全不同。

實務上的意義:鏡面架構必須量到基座運動(或量到 LOS 本身)。只在鏡面上裝一顆陀螺、把它的讀值壓到零,得到的是最糟的結果。

七、把四個決策串起來

以本系列那台 140 mm 兩軸吊艙為例,四個決策這樣落:

幾軸? 兩軸。任務是對地觀測,天底禁區在 Kt = 0.10 時只有 0.54°(第二十四篇算過),可以接受。不需要影像去旋轉,所以第三軸的重量花不下去。

直驅還是減速? 直驅。θ_max = 0.1 mrad 這條規格直接把減速排除——諧波的背隙是它的 2.9 倍起跳。

粗精複合? 不用。f_bw = 12 Hz 對單級機構不是難事,只要 f_res ≥ 60 Hz第二十五篇的分離規則)。粗精複合是給更高頻寬需求的。

動框架還是動鏡? 動框架。球型吊艙的內框慣量已經夠小(第二十三篇那個 12 mm 力臂就是球型化的成果),沒必要引入 2:1 的複雜度與基座敏感度。

四個決策全部走「簡單那一邊」,而每一個都有一個具體數字撐著。這才是架構決策該有的樣子——不是憑經驗選,是每一個選擇都能說出「因為那個數字」。

本文的分析界線

  • 背隙的量級是通用規格範圍,不是特定產品的實測。 高階諧波減速可以做到比 1 arcmin 更好,某些預壓設計的行星減速也優於 3 arcmin。要用真實型號的規格書重算。
  • 禁區那節假設「力矩夠」。 它算的是轉速約束。實際選型要兩個約束一起看,第二十四篇講的「兩側夾擊」就是這件事。
  • 粗精複合那組數字是示意,不是設計。 慣量小三個數量級、頻寬高一個數量級是我設的比例,用來說明「平方項不可怕」這個機制。真實的精級設計要從光學行程與致動器規格反推。
  • 鏡面那節只做二維、單一鏡面。 真實的鏡面架構常常是兩片鏡或鏡+框架的混合,耦合關係更複雜。2:1 這個關係本身不變,但完整的補償律要看配置。
  • 沒有處理熱、環境與壽命。 溫度會改變預壓與間隙,這對背隙與摩擦都有影響,屬於另一個題目。

想一起把這件事做完整

要幫你判斷架構,需要這些:

  • 穩定精度規格 θ_max,以及它是從哪裡來的(影像品質?定位?光通訊對準?)
  • 可接受的天底禁區半角(這一項通常沒人先想,但它直接決定軸數與 Kt
  • 需不需要影像去旋轉
  • 需要的抗擾頻寬,以及機構第一階共振落在哪
  • 體積與重量上限
  • 匯流排電壓與可用電流

沒有的就說沒有,我不會幫你猜一個值填進去。

參考文獻

  • Hilkert, J. M. (2004). A comparison of inertial line-of-sight stabilization techniques using mirrors. Proc. SPIE 5430, Acquisition, Tracking, and Pointing XVIII. DOI: 10.1117/12.541808
  • Dynamics Modeling and Theoretical Study of the Two-Axis Four-Gimbal Coarse–Fine Composite UAV Electro-Optical Pod (2020). Applied Sciences, 10(6), 1923. DOI: 10.3390/app10061923
  • Modeling and Stability Analysis of Coarse–Fine Composite Mechatronic System in UAV Multi-Gimbal Electro-Optical Pod (2020). Link
  • Ekstrand, B. (2001). Equations of Motion for a Two-Axes Gimbal System. IEEE Transactions on Aerospace and Electronic Systems, 37(3), 1083–1091. Link
  • Hilkert, J. M. (2008). Inertially Stabilized Platform Technology: Concepts and Principles. IEEE Control Systems Magazine, 28(1), 26–46.
  • Masten, M. K. (2008). Inertially Stabilized Platforms for Optical Imaging Systems. IEEE Control Systems Magazine, 28(1), 47–64.
  • Mokbel, H. F., Ying, L. Q., Roshdy, A. A., et al. (2012). Design Optimization of the Inner Gimbal for Dual Axis Inertially Stabilized Platform Using Finite Element Modal Analysis. International Journal of Modern Engineering Research.

上一篇:雲台文獻地圖。要算馬達?雲台馬達選型計算機。想追這個系列?把 LocalPapa Notes 加進書籤。

Dev Log: How Many Axes, Direct or Geared, Frame or Mirror — Four Decisions You Make First and the Literature Covers Least

This is the twenty-ninth post in the LocalPapa Notes dev-log series, and the twelfth on gimbals.

After the previous post laid the literature out as a map, one cell stood out: architecture choice has the fewest papers, and yet it is the first decision in the design flow.

The reason is not hard to see. "Two axes or three", "direct drive or geared" — the answers depend on mission requirements, which does not make a paper with general validity. So the topic drifts into vendor whitepapers and word of mouth.

But these can be computed. This post puts a number on each of the four decisions.

The literature

Academic work on architecture is thin, but not absent:

Hilkert, J. M. (2004). A comparison of inertial line-of-sight stabilization techniques using mirrors. Proc. SPIE 5430, Acquisition, Tracking, and Pointing XVIII.

The key paper on mirror stabilization. It notes that the law of reflection introduces an inherent 2:1 relationship in the axis perpendicular to the LOS, and — the most important sentence — that simply stabilizing the mirror in inertial space will not stabilize the LOS; the 2:1 in fact renders the mirror configuration particularly susceptible to base motions. It compares several techniques combining mechanisms with inertial and relative-motion sensors, discussing the trade-offs of each. Section six below unpacks that sentence.

Dynamics Modeling and Theoretical Study of the Two-Axis Four-Gimbal Coarse–Fine Composite UAV Electro-Optical Pod (2020). Applied Sciences, 10(6), 1923.

A full treatment of the coarse-fine architecture. Structure designed with finite element analysis and stress study of key components, frames in 7075 aluminium alloy to meet an under-1 kg ultralight requirement; the transmission path and kinematic coupling compensation matrix between the two-axis four-gimbal structures derived from a Euler rigid body dynamics model. Actuation uses an ultrasonic motor (USM) as the coarse stage and a voice coil motor (VCM) as the fine stage. A companion paper substitutes that model into DOB suppression control and reports disturbance rejection up to 90% better than PID and 25% better than a conventional DOB.

Ekstrand, B. (2001). IEEE TAES 37(3). The paper this series keeps returning to — its inertia symmetry condition shows that part of the cross-coupling can be designed out at the architecture stage, which is the most direct theoretical basis for putting architecture decisions before control design.

These papers model and control their respective architectures. None answers "which one should I choose" — that would require comparing architectures under one set of specs, and a paper usually covers only one. The comparisons below are my own, using the constants this series has carried throughout.

Scope of citation in this post

The method and contribution descriptions above come from abstracts and public paper pages. This environment returns 403 for most publishers, so I did not verify full texts and no unverified experimental values are quoted. The backlash magnitudes (harmonic 1–3 arcmin, planetary 3–10 arcmin) are the commonly published ranges for those transmission types, not any one vendor's data; the conversions and comparisons are my own.

1. Four decisions, and the order matters

Four architecture decisions, and the order matters Four architecture decisions, and the order matters How many axes? fixes keep-out and frame count Set by the mission: how big a nadir keep-out is acceptable, is derotation needed Direct drive or geared? fixes backlash and effective Kt Set by the stabilization spec: once backlash exceeds θ_max no controller recovers it Coarse-fine second stage? fixes the bandwidth ceiling Set by bandwidth: a second stage is only needed where one stage cannot reach Move the frame or a mirror? fixes inertia and base sensitivity Set by inertia and volume — but it costs you the 2:1 Why the order cannot be swapped: axis count sets the frame layout, the frame layout sets inertia, and only then cantransmission and actuation be discussed. Picking a motor first and then deciding the axis count means redoing all of it.
Figure 1: The order of the four architecture decisions, with what sets each one and what each one fixes.
  • How many axes? — Set by the mission: how large a nadir keep-out is acceptable, and whether image derotation is needed.
  • Direct drive or geared? — Set by the stabilization spec: once backlash exceeds θ_max, no controller recovers it.
  • A coarse-fine second stage? — Set by the bandwidth requirement: only needed where one stage cannot reach.
  • Move the frame or a mirror? — Set by inertia and volume, but it costs you the 2:1.

The reason the order cannot be swapped is practical: axis count sets the frame layout, the frame layout sets inertia, and only then can transmission and actuation be discussed. Picking a motor first and then deciding the axis count means redoing all of posts 23 and 24.

2. Axis count: the third axis does not fix gimbal lock

What each axis count actually solves — the third axis does not fix gimbal lock What each axis count actually solves — the third axis does not fix gimbal lock Two-axis (az + el) Three-axis (+ roll about LOS) Two-axis four-gimbal Pointing DOF 2 (enough for any direction) 2 (the third does not change pointing) 2 Gimbal lock Diverges near nadir Still diverges — roll is about the optical axis Outer frame takes over; no divergence Image rotation Not compensated Compensated (what the third axis is really for) Depends on layout Rejection bandwidth Limited by one mechanical stage Same as left Fine stage an order higher Mass cost Baseline + one motor and frame + fine actuator and optics A common misconception is that a third axis dodges gimbal lock. A roll axis about the optical axis does not change wherethe line of sight points, so it does not remove the singularity. Removing it requires another axis outboard, which is thetwo-axis four-gimbal layout. The third axis buys derotation and payload attitude, not pointing authority. If your problem is that the gimbal cannotkeep up near nadir, a third axis will not help.
Figure 2: What each axis count actually solves. The middle column is coloured because it is the most misunderstood.

This is the misconception this post most wants to correct.

A roll axis about the optical axis does not change where the line of sight points. It rotates the image, not the LOS direction. Therefore:

  • It does not fix gimbal lock. The nadir keep-out computed in post 24 exists on a three-axis gimbal exactly as it does on a two-axis one, because the keep-out comes from the azimuth axis's 1/cos(el) divergence and the third axis takes no part in that mapping.
  • What it does fix is image rotation. When the airframe rolls, the picture rolls with it; the third axis rolls it back. That has real value for observation and target interpretation — but it is a different requirement.

Actually removing gimbal lock requires another axis outboard — the two-axis four-gimbal layout. The extra outer frame takes over as the inner frame approaches the singularity, keeping the inner frame well away from cos(el) → 0.

So: if your problem is "it cannot keep up near nadir", a third axis will not help. That judgement is worth money, because a third axis costs a whole motor, frame, harness and power budget.

3. Direct or geared: backlash is a hard wall

Transmission: backlash is a hard wall no controller climbs Transmission: backlash is a hard wall no controller climbs 0 0.5 1 2 3 Angular error (mrad) θ_max = 0.1 mrad (this series' stabilization spec) Direct drive (zero backlash) 0 Harmonic 1–3 arcmin = 2.9–8.7 × θ_max Planetary 3–10 arcmin = 8.7–29.1 × θ_max Backlash is a mechanical dead zone, not a disturbance feedback can suppress — at the instant of reversal the feedback has no idea what happened. A harmonic drive's backlash alone is 2.9–8.7 times θ_max, so for 0.1 mrad-class stabilization,direct drive is very nearly the only option.
Figure 3: The backlash range of three transmission types, on the same ruler as this series' θ_max. The green dashed line is θ_max.

Get the units straight first. 1 arcmin = π/180/60 rad = 0.2909 mrad.

  • Harmonic drive, backlash 1–3 arcmin = 0.291–0.873 mrad
  • Planetary gearbox, backlash 3–10 arcmin = 0.873–2.909 mrad
  • Direct drive, backlash = 0 (rotor coupled straight to the load axis, torque through the magnetic field)

The stabilization spec this series has used throughout is θ_max = 0.1 mrad. Compare:

A harmonic drive's backlash alone is 2.9 to 8.7 times θ_max; a planetary gearbox's is 8.7 to 29.1 times.

What makes the comparison decisive is that backlash is not a disturbance feedback can suppress — it is a mechanical dead zone. At the instant of reversal the motor turns and the load does not, and the feedback has no idea what happened. No control law compensates something it cannot sense.

The conclusion is hard: for 0.1 mrad-class stabilization, direct drive is very nearly the only option. Conversely, if your spec is 1 mrad rather than 0.1, a harmonic drive is back on the table. Fix the spec first, then choose the transmission — not the other way round.

Note also an internal contradiction in harmonic drives: the flex spline that makes them nearly backlash-free is, by design, a part that deforms. Under load it winds up — near-zero backlash and high torsional stiffness are mutually exclusive in a harmonic drive — and torsional stiffness feeds straight into f_res, which is post 25's ceiling on the rate loop.

4. Gearing does not let you escape the nadir keep-out

The nadir keep-out depends only on effective Kt, not on the gear ratio The nadir keep-out depends only on effective Kt, not on the gear ratio 0 4 8 13 17 Effective Kt = ratio × motor Kt (N·m/A) Nadir keep-out half-angle (deg) N=1 Kt_eff=0.05 keep-out 0.27° N=5 Kt_eff=0.25 keep-out 1.34° N=20 Kt_eff=1.00 keep-out 5.36° N=50 Kt_eff=2.50 keep-out 13.52° ω_max at the load = 0.7·V_bus / (N · Kt_motor) — the ratio multiplies torque and divides speed by exactly the samefactor, so only the effective Kt survives into the keep-out formula. Gearing does not let you escape the keep-out; it just moves where you pay. You buy torque capacity and pay with a largerkeep-out, and backlash on top. That is why high-precision gimbals are nearly always direct drive — not because directdrive is better, but because the other road charges twice.
Figure 4: Keep-out half-angle against effective Kt, with four gear ratios marked. They all land on the same curve — that is the point.

This is the section I found most satisfying to work out.

Post 24 derived the nadir keep-out:

`` keep-out half-angle = asin( ω_LOS · Kt / (0.7 · V_bus) ) ``

Intuition says: add a gear ratio N, multiply torque by N, and surely you can use a low-Kt motor and shrink the keep-out?

No. Work the gearing through:

`` effective Kt at the load = N · Kt_motor ω_max at the load = (0.7 · V_bus / Kt_motor) / N = 0.7 · V_bus / (N · Kt_motor) ``

The ratio multiplies torque and divides speed by exactly the same factor, and the two cancel. Substituted into the keep-out formula, N disappears entirely — only the effective Kt survives.

With Kt_motor = 0.05, a 24 V bus and ω_LOS = 90°/s:

  • N = 1 (direct) — effective Kt 0.05, load ω_max 19251 °/s, keep-out 0.27°
  • N = 5 — effective Kt 0.25, ω_max 3850 °/s, keep-out 1.34°
  • N = 20 — effective Kt 1.00, ω_max 963 °/s, keep-out 5.36°
  • N = 50 — effective Kt 2.50, ω_max 385 °/s, keep-out 13.52°

All four land on the same curve. Gearing does not let you escape the keep-out; it just moves where you pay: you buy torque capacity and pay with a larger keep-out, plus backlash on top.

That is why high-precision gimbals are nearly always direct drive — not because direct drive is better, but because the other road charges twice.

5. Coarse-fine: why the fine stage needs less torque

When one mechanical stage cannot reach the required bandwidth, add a second. The intuitive worry is that a fine stage with ten times the bandwidth faces α = (2πf)² · θ_max, a square law — surely the torque demand explodes?

The arithmetic says otherwise. Using post 24's inner-frame inertia J = 0.0036 at f_bw = 12 Hz as the coarse stage, and a fine stage with inertia three orders smaller (a voice-coil-driven small mirror versus a whole inner frame) and bandwidth an order higher:

  • CoarseJ = 3.6×10⁻³, f = 12 Hzα = 0.568 rad/s²T = 2.05×10⁻³ N·m
  • FineJ = 3.6×10⁻⁶, f = 120 Hzα = 56.85 rad/s²T = 2.05×10⁻⁴ N·m

α is 100× larger, but inertia is 1000× smaller, and the net result is that the fine stage needs only 10% of the coarse stage's torque.

That is the physical reason coarse-fine works: the square term is frightening, but inertia is a linear term and you can make it very small. The Applied Sciences paper above uses an ultrasonic motor for the coarse stage and a voice coil motor for the fine stage — exactly this division: coarse handles travel, fine handles bandwidth.

The costs: another actuator, more optics, another layer of kinematic coupling to compensate (which is what that paper's compensation matrix is for), plus volume and mass. If one stage suffices, do not add a second.

6. Frame or mirror: the 2:1 and a trap

Mirror stabilization: the 2:1, and one counter-intuitive consequence Mirror stabilization: the 2:1, and one counter-intuitive consequence (a) Rotate the mirror by θ, the LOS moves 2θ ΔLOS mirror incoming mirror 15° → ΔLOS = 30° (twice) (b) Holding the mirror inertially still is no stabilization at all ΔLOS mirror inertially fixed (θ = 0) base rotates φ mirror still, base 20° → ΔLOS = -20° (all of it passes through) ΔLOS = 2·θ_mirror − φ_base To hold the LOS: θ_mirror = φ_base / 2 This is Hilkert 2004's point: the 2:1 of the law of reflection makes the phrase "mirror stabilization" a trap. Therequirement is a still LOS, not a still mirror — and the latter lets 100% of the base motion through. The upside is realthough: a moving mirror has far less inertia than a whole moving frame, and the same LOS rate needs only half themechanical rate.
Figure 5: Two consequences of the law of reflection. Left is the 2:1 amplification; right is the trap — hold the mirror inertially still and 100% of the base motion reaches the line of sight.

From the law of reflection (2D, math convention): outgoing direction angle = 2α + 180° − ψ, where α is the mirror normal's angle and ψ the incoming direction. Differentiating in each variable:

`` ΔLOS = 2 · θ_mirror − φ_base ``

The benefit is the first term. Rotate the mirror by θ and the LOS moves — the same LOS rate needs only half the mechanical rate. Add that a moving mirror has far less inertia than a whole moving frame, and that is the real advantage of mirror stabilization.

The trap is the second term. Suppose you "stabilize the mirror in inertial space", i.e. θ_mirror = 0. Then:

`` ΔLOS = 2 · 0 − φ_base = −φ_base ``

100% of the base motion reaches the LOS; nothing has been stabilized at all. This is precisely what Hilkert 2004 means by "simply stabilizing the mirror in inertial space will not stabilize the LOS".

Actually holding the LOS requires:

`` θ_mirror = φ_base / 2 ``

The mirror must move at half the base rate, not stay still. That is counter-intuitive enough to be worth writing on a wall — in a mirror architecture, "stabilization" means holding the LOS still, not the mirror, and the two are entirely different control objectives.

The practical implication: a mirror architecture must measure base motion (or measure the LOS itself). Putting one gyro on the mirror and driving its reading to zero produces the worst possible outcome.

7. Putting the four together

For this series' 140 mm two-axis pod, the four decisions land like this:

How many axes? Two. The mission is ground observation, and the nadir keep-out at Kt = 0.10 is only 0.54° (post 24) — acceptable. No image derotation is needed, so the third axis's mass cannot be justified.

Direct or geared? Direct. The θ_max = 0.1 mrad spec rules out gearing outright — a harmonic drive's backlash starts at 2.9× it.

Coarse-fine? No. f_bw = 12 Hz is not demanding for one stage, provided f_res ≥ 60 Hz (post 25's separation rule). Coarse-fine is for higher bandwidth requirements.

Frame or mirror? Frame. The ball pod's inner-frame inertia is already small (the 12 mm moment arm in post 23 is exactly what the ball shape buys), so there is no reason to take on the 2:1 complexity and base sensitivity.

All four land on the simpler side — and each is backed by a specific number. That is what an architecture decision should look like: not chosen from experience, but able to name the number behind every choice.

The limits of this analysis

  • The backlash figures are generic type ranges, not measurements of any product. High-grade harmonic drives can beat 1 arcmin, and some preloaded planetary designs beat 3 arcmin. Recompute from the datasheet of a real part number.
  • The keep-out section assumes torque is sufficient. It computes the speed constraint only. Real selection needs both constraints at once — the "squeezed from both sides" in post 24.
  • The coarse-fine numbers are illustrative, not a design. Three orders of inertia and one order of bandwidth are ratios I chose to show the mechanism, namely that the square term is not the problem. A real fine stage is sized backwards from optical travel and actuator specs.
  • The mirror section is 2D with a single mirror. Real mirror architectures often use two mirrors, or a mirror plus a frame, with more complex coupling. The 2:1 itself does not change, but the full compensation law depends on the layout.
  • No thermal, environmental or life considerations. Temperature changes preload and clearance, which affects both backlash and friction. That is another topic.

How to work through this with me

To help judge an architecture, I need:

  • The stabilization spec θ_max, and where it came from (image quality? geolocation? optical comms alignment?)
  • The acceptable nadir keep-out half-angle (usually nobody thinks of this first, but it drives axis count and Kt directly)
  • Whether image derotation is needed
  • The required rejection bandwidth, and where the mechanism's first resonance sits
  • Volume and mass limits
  • Bus voltage and available current

If something's missing, say it's missing — I won't guess a value and fill it in for you.

References

  • Hilkert, J. M. (2004). A comparison of inertial line-of-sight stabilization techniques using mirrors. Proc. SPIE 5430, Acquisition, Tracking, and Pointing XVIII. DOI: 10.1117/12.541808
  • Dynamics Modeling and Theoretical Study of the Two-Axis Four-Gimbal Coarse–Fine Composite UAV Electro-Optical Pod (2020). Applied Sciences, 10(6), 1923. DOI: 10.3390/app10061923
  • Modeling and Stability Analysis of Coarse–Fine Composite Mechatronic System in UAV Multi-Gimbal Electro-Optical Pod (2020). Link
  • Ekstrand, B. (2001). Equations of Motion for a Two-Axes Gimbal System. IEEE Transactions on Aerospace and Electronic Systems, 37(3), 1083–1091. Link
  • Hilkert, J. M. (2008). Inertially Stabilized Platform Technology: Concepts and Principles. IEEE Control Systems Magazine, 28(1), 26–46.
  • Masten, M. K. (2008). Inertially Stabilized Platforms for Optical Imaging Systems. IEEE Control Systems Magazine, 28(1), 47–64.
  • Mokbel, H. F., Ying, L. Q., Roshdy, A. A., et al. (2012). Design Optimization of the Inner Gimbal for Dual Axis Inertially Stabilized Platform Using Finite Element Modal Analysis. International Journal of Modern Engineering Research.

Previous post: A Map of the Gimbal Literature. Want to size a motor? Gimbal Motor Sizing Calculator. Want to follow the series? Bookmark LocalPapa Notes.

探索 61 個隱私優先的瀏覽器工具
全程本地運算、檔案不上傳。
前往 LocalPapa →
Explore 61 privacy-first browser tools
Everything runs locally — your files never leave your device.
Visit LocalPapa →

想看英文版?點右上角 EN 切換語言。

Prefer Chinese? Tap at the top-right to switch.